Until 2024 we hired the way most companies still do. A take-home or a demo project, some coding questions, a conversation about what they built. It worked because building the thing was hard. If someone shipped a clean, working application, that meant something.
It isn't hard anymore, and the consequence is mechanical rather than philosophical: every take-home we received in our last hiring rounds looked the same, because every take-home was model-written, and you can't see the difference on screen. In effect, teams still relying on them are only filtering for candidates willing to spend an evening supervising an agent.
LeetCode-style tests have the opposite problem. They measure something real, but what they measure is pattern recall under a clock, and that was never a great predictor for senior work even before the tools changed. You won't learn how a senior engineer handles a schema change by watching them invert a binary tree. Worse, the recall those tests reward is precisely the work agents absorbed. You'd be selecting hardest for the skill that matters least.
What broke, in one sentence: hiring processes were built to evaluate output, and output stopped carrying information about the person. Neither format shows the fundamentals anymore, and neither shows the agent skills. So we replaced both with two working sessions in a real codebase, one without agents and one with, where we watch how people actually work.
Neither session stands alone. They're stages three and four of a longer process: CV screening against a rubric and a thirty-minute intro call come before them, a culture interview and references after, because when most of the code is written by agents, how someone shares context and works with the people around them matters at least as much as their technical skills.
The first session tests the fundamentals, with no AI in it. The candidate gets a codebase they haven't seen, a handful of real tickets, one of them a production bug, and sixty minutes. It's a normal working session, not a test: they pick up the tickets and get on with it, and we watch. Writing code has become easy, but understanding how a system fits together hasn't, and if the fundamentals aren't there, the candidate has nothing to validate an agent's output against. What good looks like is specific: reading the data flow and failure modes before touching code, putting the production bug first and being able to say what it affects, finding the root cause rather than the symptom, adding the test that would have caught the bug. This round catches people who've been relaying an LLM through a few years of remote work, which is more common than the industry likes to admit. It's also where we find out whether they can talk through a system's design fluently in English, which we screen harder on than most.
The second session is the same shape of work with the opposite constraint: a fresh set of tickets, and whatever AI coding tools the candidate would normally use. Now we watch what they do when an agent is doing most of it, scored on six criteria: whether they verify what the agent produces, how they orchestrate the tools, whether they break work into pieces before generating, how they set up prompt and context, how they debug, and whether they keep track of everything they've started. Whether they plan first, whether they read what comes back, and whether they can explain a change they didn't type.
This second session is where a recent candidate with thirteen years of experience fell apart. His CV looked good and the fundamentals session went well. Then he was allowed to use agents, and from that point on he stopped reading. Every change got accepted, every prompt got a "go ahead", and he closed tickets without checking any of them. Most companies would have hired him. We would have too, two years ago, because most processes never see the hour we saw. If you're still filtering on take-homes and coding tests, he's getting through your interviews right now, and you'll find out when his code reaches review.
If you're rebuilding your own process, the two-session structure is the principle worth stealing, more than any specific task. Fundamentals without the agent, because they're what make verification possible. Real work with the agent, because that's the job. The parts of hiring that were never about the artefact, the screening, the culture conversation, the references, keep their old role; if anything, the reference check matters more now.
And if you're evaluating a vendor rather than a candidate, the same logic applies with one addition: ask to see the process. Any firm placing engineers should be able to show you what they test, how it's scored and why people fail. We publish ours in full. A vetting process that can't survive publication was measuring the wrong things.
The full process, with both scorecards and what good and bad look like in each session, is in The Next 10X Engineer. And if you want to know how your current hiring compares with other leaders', the Engineering Leader Benchmark is open.