Writing
Vibe Coding, One Year Later: Agentic Workflows and Engineered Loops
What changed when an agentic loop inside an editor became asynchronous, context-managed work, and why engineering the loop now matters as much as the model.
When I first wrote about vibe coding in April 2025, I’d been working that way for about three months, roughly as long as the phrase had existed. I wanted to name both the promise and the catch: this genuinely worked, and you still had to know how to engineer software. A personal toy you can just pump out. But the moment you want to share it, the moment it stops being private, that’s where the trouble starts.
Karpathy had coined the phrase only that February; my first write-up came with a graphic, 15 Rules of Vibe Coding.
A year later, whether AI can help write code isn’t the question. That argument is over. What matters is what comes after the first burst of productivity: after the prompt, the generated code, and the thrill of watching an application appear.
One shift made this possible and one made it matter: agentic workflows made delegation real, and loop engineering is what makes it repeatable.
Cursor Agent was already agentic in an important sense. It could inspect a repository, edit files, run commands, and work through a problem over multiple tool calls. But the workflow I described still revolved around me driving one session-scoped agent loop. I supplied the context, watched the run, noticed when it wandered, and decided what to paste back into the chat.
What I mean by an agentic workflow now is a larger unit of delegation. It packages a bounded objective with the context, tools, permissions, state, and output it needs. That encapsulation matters because the work no longer has to stay inside an active back-and-forth with me. I can delegate it, let it run asynchronously, and come back to an artifact, a result, and a record of what happened. Claude Code’s parallel-agent model and the Codex app are concrete examples of that shift.
Loop engineering is the complementary advance. It defines how those agentic units are checked, connected, repeated, redirected, or stopped. The early Cursor workflow put an agentic loop inside the editor. Agentic workflows give a piece of work its own context and lifecycle; loop engineering governs what happens within and across those lifecycles.
Why not a year ago?
If the agent was already there, why wasn’t I engineering loops around it a year ago? Two reasons, and neither was the model. The tools were single-session and synchronous: Cursor couldn’t run a background or parallel agent until version 0.50, which shipped less than two weeks after that first post. There were no separate agentic contexts to orchestrate: there was one chat, and me watching it. And the economics pointed the same way. Pro billed a fixed pool of requests each month; every speculative run drew down a budget I could feel, so you kept the one agent you had on a short, supervised leash rather than spinning up five to explore in parallel.
Both constraints moved. Usage now meters on predictable, resetting windows: Claude Code’s Pro and Max plans refill on a five-hour session cycle and a weekly one, so long-running work can stop at a limit and resume on schedule instead of dying at a hard cap. And a background or parallel agent, each with its own context, is now just how the tools work. Long-running, asynchronous, many-agent work stopped being reckless and became ordinary. That, as much as any model improvement, is what made loop engineering worth doing.
The loop moved up a level
Once those constraints lifted, the loop could rise above the session. A year ago the agent was real but the harness around it was me: the 15 rules (start from a template, open a fresh chat, paste the error back, test locally, commit before getting too far) describe a loop I was running by hand, as its memory, scheduler, evaluator, and recovery mechanism. The Model Context Protocol was already widening what the agent could reach, but more tools were never the bottleneck; orchestration was. What changed is that the task, not the session, became the unit I delegate. And running one asynchronously is not the same as firing it and forgetting it. It needs durable state, checkpoints, and a trace of what the agent decided, or background execution is just unattended execution.
From feedback habits to loop engineering
Several of those early instincts still matter (isolate the task, expose failures, preserve a safe point, test before proceeding, release carefully), and I later reduced them to the one I cared about most: test before proceeding. What’s changed is that I no longer want to run every turn of that cycle by hand. Once work travels in bounded units, the problem shifts from writing a better prompt to deciding what happens before, during, and after each run:
- Define the task and the evidence that would count as success.
- Assemble only the relevant context and tools.
- Let the agent work inside its harness.
- Run deterministic checks wherever possible.
- Evaluate the result against the original intent, not just whether the code compiled.
- Accept it, send back targeted feedback, try a different approach, or stop for a person.
Put safety and deployment inside the loop
Securing secrets and deploying early were already instincts I had, but neither goes far enough for an agentic workflow. An agent should receive only the credentials it needs, at the narrowest useful permissions; risky tools should be sandboxed or approval-gated; and inputs, outputs, and tool calls may need validating at their actual boundaries, not just at the beginning and end of a long run.
Deploying early still matters (the real environment reveals problems local execution cannot), but “deploy” should be a controlled transition in the loop: build, check, release, smoke-test, observe, and keep a rollback path. A successful command is not a successful deployment. After release, tracing and production feedback belong in the next pass through the loop, alongside the tests that confirm known behavior still works.
When one loop is not enough
The outer loop becomes even more useful when one agent should not make every judgment. Instead of stretching one prompt across planning, implementation, review, security, and evaluation, I can make those responsibilities explicit as nodes in a graph. State moves between them, and conditional edges decide what runs next, an approach reflected in modern graph-oriented agent workflows.
One concrete example is a strategy-idea screen I built for a trading-research side project: before an idea consumes a real coding-and-backtest cycle (the expensive part of the pipeline), it goes through a Graph-of-Agents pattern, adapted from Yun and colleagues’ 2026 framework, that decides whether it deserves the investment.
The graph has four critic agents, each looking for a different way an idea can fail. One asks whether there is a plausible causal mechanism. Another checks whether something substantially similar has already been tried. A third asks whether the likely edge could survive the cost model for this round. The fourth checks whether the idea is concrete enough to implement without the coder having to invent missing behavior.
Not every idea gets all four critics. An orchestrating agent (not me, running each round by hand) samples the two or three that matter for that kind of strategy, then calls each selected critic once for the whole batch rather than once per idea. Each critic’s own confidence in its verdict becomes its weight in the pool: an approximation of the paper’s full peer-scoring, chosen to keep triage cheap relative to the backtest cycle it protects.
If the selected critics agree, the triage stops there. There is no point paying for another round of conversation just so the agents can congratulate one another. If they disagree, the orchestrator runs one bounded exchange: the higher-confidence critic’s verdict goes to the lower-confidence one, which gets a chance to revise. If it genuinely changes its mind, the revision goes back for one final check. And if the critics still can’t converge, that unresolved disagreement, not every run, is what gets surfaced. By default it comes to me, but it doesn’t have to: I could route it to a Fable-model agent as arbiter instead and reserve myself for the disagreements sticky enough to actually need a person. That gives useful message passing without letting the review turn into an open-ended debate, and keeps a person in the loop only for the judgment calls that actually need one.
The composite verdict (clear, caution, or dead) comes from a weighted average of the critics’ judgments, mean-pooled by default and max-pooled only when one critic’s objection should clearly dominate. Because a false dead verdict would quietly kill an idea before it ever reached a real backtest, the first live round ran in shadow mode: every verdict was computed and logged, but nothing was pruned. Only after comparing those shadow calls against the real backtest results did I trust the threshold enough to let it filter ideas for real.
That is a long way from pasting an error into a chat window. The model is no longer just generating code. Multiple asynchronous, context-encapsulated agents are producing and challenging judgments inside a loop whose routing, cost, stopping conditions, and risk controls were designed in advance.
What changed for me
A year ago, the work was about operating one agent loop well. Now it’s about what surrounds that loop: what context survives, what can run on its own, what evidence has to come back, and what happens next. The human moved up a level too: I spend less time asking “what should I type next?” and more deciding which judgments the system can make on its own, and where a person still has to be accountable for the result.
It was never really a jump from vibe coding to agents; the agent was already there. The jump was from working inside a single session to engineering the loops around many of them, so a burst of generated code becomes work you can check, route, repeat, and trust.
The vibe is still a great place to begin. It’s just no longer where the engineering ends.