We've published our vetting standard. Both assessments, both scorecards, what good looks like, what bad looks like: all of it, openly. Before you read it, this is the thinking behind the six criteria we score when the candidate has an agent open, and why the session is built the way it is.
The AI-native assessment is the second of two working sessions in a real codebase. By the time a candidate reaches it, they've passed a fundamentals session in a codebase they'd never seen, with real tickets and no AI tools, because if you can't reason about a system without the agent, you can't tell when the agent is wrong. Then the second session starts: a fresh set of tickets, and whatever AI coding tools they'd normally use. An agent doing most of the production, and us watching what the human actually contributes.
The first criterion is AI output verification. Good looks like catching hallucinated APIs, subtly wrong logic or security gaps in generated code before it ships. Bad looks like trusting it because it compiles. The failure has a texture you learn to spot within minutes: the candidate stops being an engineer and becomes a relay between the model and the codebase. Every change accepted, every prompt answered with "go ahead", tickets closed without checking any of them. It's exactly how the thirteen-year developer in the paper's introduction came apart, and it's invisible to any process that only looks at what he shipped, because what he shipped compiled.
The second is tool orchestration, and it's where the multiplier lives. Good looks like knowing when to hand work to the agent, the scaffolding, the boilerplate, the tests, and when to hand-code the novel logic and the awkward edge cases, then running the two in parallel in worktrees so agents don't collide. Bad looks like one window used for everything, or its mirror image, hand-coding things an agent does better out of pride. Almost no interview process anywhere tests this, which is remarkable given it's the difference between an engineer and the several-times-faster version of the same engineer.
The remaining four complete the picture. Spec-driven development: breaking a vague ticket into pieces small enough for the tool to execute well, rather than throwing the whole ambiguous ask at it in one go. Prompt and context engineering: setting up the task so the tool produces usable output in one or two passes, instead of re-prompting the same vague request in a loop. AI-accelerated debugging: using the agent for root-cause analysis, log triage and stack traces, without skipping the step of reproducing and understanding the failure yourself. And context-switch hygiene: knowing what each agent is doing, coming back to the right one, and not letting notifications drive the switching. That last one is scored in both sessions, because it's about the engineer rather than the tools, and it's the one that decides whether orchestration is a multiplier or just several wrong things happening at once.
The paper says plainly that the scorecard isn't the whole picture. Three things matter as much and don't fit on a card, so we look for them across both assessments. Judgement by risk: a copy change, a low-risk refactor and a payment migration shouldn't get the same review, and with agents this matters more, because the tool produces all three at the same speed and with the same confidence. Ownership: if you merge it, you own it, and "the agent did it" isn't an explanation we accept in an assessment or one your team will accept in an incident review. And leverage: the best engineers finish the ticket and also leave the repo better set up for the next agent, a reusable command, a tighter context file, a test that catches the class of bug rather than the instance.
Why publish the whole thing, scoring included? There are three reasons. First, the industry needs a reference point; "we vet for AI skills" has already gone the way of "senior" on a CV, and a published standard is harder to imitate than a claim. Second, candidates deserve to know what they're walking into; the engineers we want are the ones who read the standard and think "finally, someone's testing the actual job". Third, we can afford to. The scorecards aren't the moat. The moat is that the people running the sessions build production software with these tools every day, which is the only way to know what a wrong answer looks like. A staffing firm could copy our rubric tomorrow and still couldn't score it.
There's a buyer's use for the publication too, and we mean it as an invitation rather than a jab. If you're paying anyone for engineers, hold their vetting against a published standard, ours or a better one. Ask what's tested, how it's scored, what fails people. The firms doing this properly will answer in detail and probably enjoy the question. The rest will send you a slide about their talent pool.
The whitepaper walks through both sessions with the scorecards we fill in during them. If you read one chapter, read the one on the AI-native assessment. It's where the next bad hire is hiding.
Read the standard at how we vet. Download The Next 10X Engineer.