Agentic coding is a mode of software development in which an AI agent takes a goal, breaks it into steps, writes code, executes it, reads the output, and iterates, without a human directing each individual action. The engineer sets the objective and reviews the result; the agent handles the intermediate steps.
How it differs from assisted coding
Most AI coding tools respond to a single prompt and return a single completion. The engineer decides what to do next. Agentic coding inverts that: the agent decides what to do next, calling tools, running tests, reading error messages, and adjusting its approach in a loop. The human re-enters when the loop ends or when something goes wrong.
This is sometimes called a coding agent, and it relies on the model having access to a set of tools: a code executor, a file system, a terminal, sometimes a browser or an external API.
What the loop looks like in practice
A typical agentic coding cycle runs roughly as follows:
- The engineer provides a goal: "Add rate limiting to the payments endpoint, with tests."
- The agent reads the relevant files to understand the existing code.
- It writes a plan, then writes the code.
- It runs the tests. If they fail, it reads the error and tries again.
- It returns a diff or a summary for the engineer to review.
The number of steps between human inputs can be anywhere from three or four to several dozen, depending on the task and the tool.
Where the context window matters
Agentic loops accumulate context quickly. Each tool call, each error message, each file read adds tokens. A long-running agent on a complex codebase can exhaust a model's context window, at which point earlier information falls out and the agent loses track of what it has already done. Managing what goes into context at each step, and what gets summarised or discarded, is a practical engineering concern, not an afterthought.
What can go wrong
The main risks are not dramatic. They are quieter:
- The agent confidently completes a task that was not quite what was intended, because the spec was ambiguous.
- It introduces a change in one file that breaks an assumption in another file it did not read.
- It passes its own tests because it wrote both the code and the tests to the same misunderstanding.
- It loops on a failing test, making increasingly speculative changes, and produces a diff that is harder to review than the original problem.
None of these are unique to agentic coding, but the autonomous loop means mistakes can compound before a human sees them. This is why reviewing AI-generated code and maintaining a human in the loop at meaningful checkpoints matters more, not less, as agents become more capable.
The relationship to vibe coding
Vibe coding is sometimes confused with agentic coding. The distinction is about intent, not tooling. Agentic coding is a technical pattern: an agent with tool access operating across multiple steps. Vibe coding describes a working style where the engineer accepts output with minimal scrutiny. You can use an agentic tool carefully or carelessly. The tool does not determine the standard of work.
What we test for
Agentic coding changes what a vetting process needs to measure. Our how we vet process runs two sessions: a fundamentals assessment without AI tools and an AI-native assessment with them. The second session is where agentic work shows up directly. We score for AI output verification, because generated code that compiles is not the same as code that is correct, and for judgement by risk, because an agent produces a payment migration and a copy change at the same speed and with the same confidence. An engineer needs to treat them differently.
Short answers
Is agentic coding the same as using GitHub Copilot?
No. Copilot and similar tools complete single prompts and return control immediately. Agentic coding means the AI takes multiple actions in sequence, running code and reading results, before returning to the engineer.
Do you need a special tool to do agentic coding?
Yes, in practice. The agent needs access to tools: a code executor, file system, terminal or similar. Standard chat-based AI models without tool access cannot run agentic loops. Several products now support this natively.
Is agentic coding suitable for production work?
It can be, with appropriate checkpoints. The risk is that errors compound across steps before a human reviews them. Engineers need to scope tasks carefully, review diffs critically, and not treat a passing test suite as sufficient validation.