A coding agent is an AI system that can write, edit, run and debug code across multiple steps without a human directing each action. It operates in a loop: generating code, executing it, reading the output, and revising until a goal is met or it gets stuck.
How it differs from a code-completion tool
Most developers have used AI to autocomplete a line or generate a function on request. That is a single-turn interaction: you prompt, the model responds, you decide what to keep.
A coding agent goes further. It holds a goal, breaks it into steps, uses tools (a shell, a file system, a browser, an API), and works through those steps over many iterations. It can install dependencies, run tests, read error messages and try again. The human may only be involved at the start and the end.
This sits closer to agentic coding than to assisted writing.
What tools a coding agent typically uses
- File read and write (creating, editing, deleting files in a codebase)
- Shell execution (running commands, scripts, test suites)
- Web search or documentation lookup
- External APIs or services
- Sometimes: spawning sub-agents for parallel tasks
The agent decides which tools to call based on the current state of its task.
What it is actually doing
Under the hood, the agent is making repeated calls to a language model. Each call includes the original goal, the history of what has happened so far, and the current context: file contents, error output, test results. This accumulated context is what allows it to reason across steps rather than starting fresh each time.
The size and structure of that context matters. A poorly managed context window can cause the agent to lose track of earlier decisions or repeat work it has already done.
Where coding agents are used in practice
Common uses include:
- Writing and running a new feature end to end on a defined ticket
- Refactoring a module and verifying tests still pass
- Debugging a failing CI pipeline
- Generating boilerplate for a new service
- Triaging and fixing a reported bug with a reproduction script
More ambitious uses involve longer-running tasks, larger codebases, or agents that hand off work to one another. These introduce more risk because errors compound across steps and may be harder to catch.
The main practical risks
Coding agents can produce output that looks correct but is not. They may pass their own tests while missing edge cases. They can make changes across many files, some of which interact in ways the agent did not account for. And because the work happens quickly and without narration, a reviewer seeing only the final diff may have little visibility into what reasoning produced it.
This is why reviewing AI-generated code requires different habits from reviewing code a colleague wrote step by step. The diff may be large, internally consistent and confidently wrong in a specific area.
Human oversight at defined points, rather than only at the end, reduces this risk. See human-in-the-loop engineering for more on how teams structure that.
What we test for
Coding agents change what we look for when vetting engineers. In Session 1, candidates work without AI tools, so we can assess fundamentals like debugging and prioritisation directly. In Session 2, they work with their usual AI tools, and we score things like AI output verification and tool orchestration: whether they catch plausible but wrong generated code, and whether they know what to hand to an agent versus what to write themselves. Full details are at how we vet.
Short answers
Is a coding agent the same as GitHub Copilot?
No. Copilot and similar tools complete or generate code in response to a prompt. A coding agent works autonomously over multiple steps, running code, reading output and revising without a human directing each action.
Can a coding agent work on an existing codebase?
Yes, though effectiveness depends on codebase size, documentation quality and how well the agent can navigate unfamiliar structure. Larger or poorly documented codebases increase the chance of the agent making changes that conflict with existing conventions.
Do coding agents replace engineers?
Not at present. They handle bounded, well-specified tasks reasonably well but require an engineer to set goals, review output and intervene when the agent goes wrong. The skill shifts; the need for judgement does not.