What we score with AI tools off

Hiring3 min read

The most contrarian hour in our hiring process has no AI in it at all. A codebase the candidate has never seen, a handful of real tickets, one of them a production bug, sixty minutes with a tech lead, and no agents allowed. Candidates are regularly surprised by it. A few have asked whether we're behind the times.

The opposite. The fundamentals session exists because of the agents, not despite them.

Here's the reasoning, and it's the sentence the whole whitepaper hangs on: if you can't reason about the system without the agent, you can't tell when the agent is wrong. On our builds, 80 to 90 percent of production code is written by a model, which means an engineer's real contribution is everything around the generation: knowing what to ask for, reading what comes back, catching the confident mistake. All of that runs on fundamentals. Writing code became easy; understanding how a system fits together didn't. So before we watch anyone work with an agent, we take the agent away and see what's left.

The session is deliberately mundane. No puzzles, no whiteboard, no binary trees. The candidate picks up the tickets and gets on with it, and we watch how they work, scoring six things.

Technical understanding comes first and shows itself fastest: depth on the stack, on the application's complex flows, and in the terminal. The strongest candidates start the application and inspect the codebase before changing anything. The weakest never start the application at all, which sounds impossible until you've watched it happen, and edit code they haven't read.

Debugging is the production bug. Good looks like working from symptom to cause in a structured way and testing the fix before calling it done. Bad looks like trial and error until something stops erroring, then closing the ticket fast and leaving it broken. This one transfers directly to the agent era, because debugging generated code you didn't write is now most of the job.

Prioritisation is quietly diagnostic. The tickets come with stated urgency, and the production bug comes first, every time. An engineer who works the interesting ticket before the burning one is telling you how they'll behave on your roadmap.

Communication is scored on whether they narrate trade-offs and say what they didn't get to, rather than going quiet, and whether they ask when a ticket is ambiguous instead of guessing. Going quiet when stuck is one of the clearest bad signals we have. English level sits alongside it: understood and understandable, spoken and written, enough to argue a point under a follow-up question rather than only recite one. When most code is agent-written, how someone shares context matters at least as much as their technical skill, which is also why a culture interview and a reference check follow the technical sessions.

And context-switch hygiene: keeping track of which tickets are open and where each one is, picking a task back up without re-reading everything. We score this again in the second session with agents running, because it's about the engineer rather than the tools, and orchestrating four agents with poor task-tracking just produces four confused workstreams instead of one.

Notice what this scorecard is really measuring. Nothing on it is new. Every item would have described a good engineer in 2015. What changed is what the items are for: they used to be the job, and now they're the licence to do the job, the thing that makes an engineer's verification of agent output worth anything at all. That's why the fundamentals session comes before the AI-native one and why failing it ends the process. There's nothing for the second session to measure.

This closes six weeks of publishing built on The Next 10X Engineer: what the job is now, what AI-native means, why the old filters broke, the AI-native scorecard, and now the fundamentals underneath it all. The paper has both sessions and both scorecards in full, and the vetting standard is public.

So a closing question, in the spirit of the opening one. Your interview process almost certainly still tests fundamentals. But if an agent sat down next to your last hire on day one, would your process have told you whether they can check its work, or only whether they could have done the work themselves in 2021?

Download The Next 10X Engineer for both assessments and both scorecards. The Engineering Leader Benchmark is open, and contributors read the report first.

Let's talk

Get a shortlist within five working days

You share the roles and the stack in a short form or a thirty-minute call. Within five working days you get named senior engineers to review, each with both scorecards.

Reviewed onClutch4.9 out of 5 from 36 reviews
ISO 27001
Certified

Book thirty minutes with Dale

The calendar is provided by HubSpot, which sets its own cookies. Load it here, or book on HubSpot's page.

Open booking page