How we vet
Every engineer we place passes two sixty-minute live assessments in a real codebase, one with AI tools off and one with their own agents running. You see both scorecards before you meet them.
-
01
Step 1
Application review
Shipped code with a commit history that shows iteration, matched to the role and stack. Not tutorial projects, and not AI-tooling claims with no code behind them.
-
02
Step 2
Intro call
Answers the question asked, not the one they prepared for.
-
03
Step 3 · Assessment 1
Fundamentals assessment
Sixty minutes in an unfamiliar codebase with AI tools switched off. It shows whether they can read a system, find the real problem and reason about trade-offs on their own.
See the scorecard -
04
Step 4 · Assessment 2
AI-native assessment
Sixty minutes on fresh tickets with their own agents running. It shows how they plan the work, check what the agent produces and decide what's safe to ship.
See the scorecard -
05
Step 5
Leadership interview and references
For senior and client-facing roles: a consistent story across two interviewers, and former managers who confirm it.
-
06
Step 6
Decision
Every scorecard considered together, with a written rationale. One interviewer's strong impression doesn't override a scorecard that didn't clear the bar.
-
07
Step 7
Selected for a shortlist
Engineers who clear every step are matched to your roles and stack, and you get their scorecards before you meet them.
Get a shortlist
Why we replaced take-homes and coding tests
Every take-home looks the same now, because every take-home is model-written. LeetCode measures pattern recall under a clock. Neither shows the fundamentals, and neither shows the agent skills.
So we replaced both with two live assessments in a real codebase: one without agents, one with. These are the scorecards we fill in for every engineer we place.
Assessment 1 · AI tools off
The fundamentals assessment
The engineer gets a codebase they haven't seen, a handful of tickets including one production bug, and sixty minutes to work through them. It runs like a normal working session, with no AI tools allowed.
See the scorecardAssessment 2 · AI tools on
The AI-native assessment
The engineer gets a fresh set of tickets in the same kind of codebase, plus whatever AI coding tools they'd normally use. The assessment shows what they do when an agent is doing most of the work: whether they plan first, whether they read what comes back, and whether they can explain a change they didn't type.
See the scorecardWhy fundamentals still matter with agents
The engineer gets a codebase they haven't seen, a handful of tickets including one production bug, and sixty minutes to work through them. It runs like a normal working session, with no AI tools allowed.
What good looks like
- Reads the data flow and failure modes before touching code
- Puts the production bug first and can say what it affects
- Traces the symptom back to its root cause
- Adds the test that would have caught the bug
What bad looks like
- Fixes the symptom and moves on
- Treats every ticket as the same risk
- Can say that a change works, but not why
- Changes shared code without checking who depends on it
The scorecard · six criteria
Communication
Narrates trade-offs and says what they didn't get to, rather than going quiet. Asks when a ticket is ambiguous instead of guessing.
English level
Is understood and understandable, spoken and written. Can argue a point and hold it under a follow-up question.
Prioritisation
The order they work the tickets makes sense given the stated urgency. The production bug comes first, every time.
Debugging
Finds root causes in a structured way, symptom to cause, rather than by trial and error. Tests the fix before calling it done.
Context-switch hygiene
Keeps track of which tickets are open and where each one is when they switch. Picks a task back up without re-reading everything.
Technical understanding
Shows depth on the stack, on the application's complex flows and in the terminal. Starts the application and inspects the codebase before changing anything.
Agents build whatever you give them
The engineer gets a fresh set of tickets in the same kind of codebase, plus whatever AI coding tools they'd normally use. The assessment shows what they do when an agent is doing most of the work: whether they plan first, whether they read what comes back, and whether they can explain a change they didn't type.
What good looks like
- Writes down assumptions, constraints and edge cases
- Flags any changes that involve data, auth or billing
- Runs the diff and asks the agent why it chose that approach
- Leaves something reusable behind
What bad looks like
- Same review depth for a copy change and a migration
- “The agent did it” as an explanation
- Re-prompts instead of fixing the context
- Accepts a plan that duplicates what's already in the repo
The scorecard · six criteria
AI output verification
Catches hallucinated APIs, subtly wrong logic or security gaps in generated code before it ships, rather than trusting it because it compiles.
Tool orchestration
Knows when to hand work to the agent (scaffolding, boilerplate, tests) and when to hand-code (novel logic, awkward edge cases), and runs the two in parallel in worktrees.
Spec-driven development
Breaks a vague ticket into pieces small enough for the tool to execute well, rather than throwing the whole ambiguous ask at it in one go.
Prompt and context engineering
Sets up the prompt and the task context so the tool produces usable output in one or two passes, instead of re-prompting the same vague ask.
AI-accelerated debugging
Uses the agent for root-cause analysis, log triage and stack traces, without skipping the step of reproducing and understanding the failure themselves.
Context-switch hygiene
Context-switch hygiene is scored again here, with agents running. A good engineer knows what each agent is doing, comes back to the right one, and doesn't let notifications drive the switching.
What separates senior engineers from fast ones
They don't fit on a scorecard, so we look for them across both assessments.
01
Judgement by risk
A copy change, a low-risk refactor and a payment migration shouldn't get the same review. Good engineers speed up on the first and slow right down on the third. Agents produce all three at the same speed and with the same confidence.
In the assessment
Does the candidate read the migration diff line by line and skim the copy change? Do they bring up blast radius before we do?
02
Ownership
If you merge it, you own it. Tests, observability, rollout and the debugging after release are part of the work whether a person or an agent wrote the code.
In the assessment
“The agent did it” isn't an explanation we accept in an assessment, and it won't be one your team accepts in an incident review.
03
Leverage
The best engineers finish the ticket and leave the repo better set up for the next agent: a reusable command, a tighter CLAUDE.md, a test that catches the class of bug rather than the instance.
In the assessment
Did anything they built during the hour make the next ticket easier, or did they leave the repo exactly as they found it?
Do you publish your assessment?
Yes. Both assessments, what good and bad look like, and the twelve criteria on the two scorecards are on this page and in The Next 10X Engineer whitepaper.
Can I see a candidate's results?
Yes. Every shortlisted engineer comes with both scorecards filled in. You see them before you meet anyone.
Who runs the assessments?
Senior engineers who build production software with these tools every day run both assessments. Recruiters don't score candidates.
Why not a take-home or a coding test?
Every take-home is model-written now, and LeetCode-style tests measure pattern recall under a clock. Neither shows the fundamentals or the agent skills.
See the scorecards before you meet anyone
Tell us the role. We send named engineers with both scorecards filled in within five working days.
Certified