Engineers reviewing a database schema together

How we vet

Every engineer we place passes two sixty-minute live assessments in a real codebase, one with AI tools off and one with their own agents running. You see both scorecards before you meet them.

The full vetting process
  1. 01

    Step 1

    Application review

    Shipped code with a commit history that shows iteration, matched to the role and stack. Not tutorial projects, and not AI-tooling claims with no code behind them.

  2. 02

    Step 2

    Intro call

    Answers the question asked, not the one they prepared for.

  3. 03

    Step 3 · Assessment 1

    Fundamentals assessment

    Sixty minutes in an unfamiliar codebase with AI tools switched off. It shows whether they can read a system, find the real problem and reason about trade-offs on their own.

    See the scorecard
  4. 04

    Step 4 · Assessment 2

    AI-native assessment

    Sixty minutes on fresh tickets with their own agents running. It shows how they plan the work, check what the agent produces and decide what's safe to ship.

    See the scorecard
  5. 05

    Step 5

    Leadership interview and references

    For senior and client-facing roles: a consistent story across two interviewers, and former managers who confirm it.

  6. 06

    Step 6

    Decision

    Every scorecard considered together, with a written rationale. One interviewer's strong impression doesn't override a scorecard that didn't clear the bar.

  7. 07

    Step 7

    Selected for a shortlist

    Engineers who clear every step are matched to your roles and stack, and you get their scorecards before you meet them.

    Get a shortlist
Assessment 1 · The fundamentals assessment

Why fundamentals still matter with agents

The engineer gets a codebase they haven't seen, a handful of tickets including one production bug, and sixty minutes to work through them. It runs like a normal working session, with no AI tools allowed.

What good looks like

  • Reads the data flow and failure modes before touching code
  • Puts the production bug first and can say what it affects
  • Traces the symptom back to its root cause
  • Adds the test that would have caught the bug

What bad looks like

  • Fixes the symptom and moves on
  • Treats every ticket as the same risk
  • Can say that a change works, but not why
  • Changes shared code without checking who depends on it

The scorecard · six criteria

01

Communication

Narrates trade-offs and says what they didn't get to, rather than going quiet. Asks when a ticket is ambiguous instead of guessing.

02

English level

Is understood and understandable, spoken and written. Can argue a point and hold it under a follow-up question.

03

Prioritisation

The order they work the tickets makes sense given the stated urgency. The production bug comes first, every time.

04

Debugging

Finds root causes in a structured way, symptom to cause, rather than by trial and error. Tests the fix before calling it done.

05

Context-switch hygiene

Keeps track of which tickets are open and where each one is when they switch. Picks a task back up without re-reading everything.

06

Technical understanding

Shows depth on the stack, on the application's complex flows and in the terminal. Starts the application and inspects the codebase before changing anything.

Assessment 2 · The AI-native assessment

Agents build whatever you give them

The engineer gets a fresh set of tickets in the same kind of codebase, plus whatever AI coding tools they'd normally use. The assessment shows what they do when an agent is doing most of the work: whether they plan first, whether they read what comes back, and whether they can explain a change they didn't type.

What good looks like

  • Writes down assumptions, constraints and edge cases
  • Flags any changes that involve data, auth or billing
  • Runs the diff and asks the agent why it chose that approach
  • Leaves something reusable behind

What bad looks like

  • Same review depth for a copy change and a migration
  • “The agent did it” as an explanation
  • Re-prompts instead of fixing the context
  • Accepts a plan that duplicates what's already in the repo

The scorecard · six criteria

01

AI output verification

Catches hallucinated APIs, subtly wrong logic or security gaps in generated code before it ships, rather than trusting it because it compiles.

02

Tool orchestration

Knows when to hand work to the agent (scaffolding, boilerplate, tests) and when to hand-code (novel logic, awkward edge cases), and runs the two in parallel in worktrees.

03

Spec-driven development

Breaks a vague ticket into pieces small enough for the tool to execute well, rather than throwing the whole ambiguous ask at it in one go.

04

Prompt and context engineering

Sets up the prompt and the task context so the tool produces usable output in one or two passes, instead of re-prompting the same vague ask.

05

AI-accelerated debugging

Uses the agent for root-cause analysis, log triage and stack traces, without skipping the step of reproducing and understanding the failure themselves.

06

Context-switch hygiene

Context-switch hygiene is scored again here, with agents running. A good engineer knows what each agent is doing, comes back to the right one, and doesn't let notifications drive the switching.

What doesn't fit on a card

What separates senior engineers from fast ones

They don't fit on a scorecard, so we look for them across both assessments.

01

Judgement by risk

A copy change, a low-risk refactor and a payment migration shouldn't get the same review. Good engineers speed up on the first and slow right down on the third. Agents produce all three at the same speed and with the same confidence.

In the assessment

Does the candidate read the migration diff line by line and skim the copy change? Do they bring up blast radius before we do?

02

Ownership

If you merge it, you own it. Tests, observability, rollout and the debugging after release are part of the work whether a person or an agent wrote the code.

In the assessment

“The agent did it” isn't an explanation we accept in an assessment, and it won't be one your team accepts in an incident review.

03

Leverage

The best engineers finish the ticket and leave the repo better set up for the next agent: a reusable command, a tighter CLAUDE.md, a test that catches the class of bug rather than the instance.

In the assessment

Did anything they built during the hour make the next ticket easier, or did they leave the repo exactly as they found it?

The Next 10X Engineer whitepaper cover

Whitepaper

The Next 10X Engineer

Why the best engineers now write less code, and how to tell them apart from everyone else.

Fourteen pages by Bo Wesdorp, our CTO: what the job is now, what AI-native means, why the old hiring process stopped working, and both scorecards in full.

FAQs

Frequently asked questions

If yours isn't here, we're happy to help.

Contact us
Do you publish your assessment?

Yes. Both assessments, what good and bad look like, and the twelve criteria on the two scorecards are on this page and in The Next 10X Engineer whitepaper.

Can I see a candidate's results?

Yes. Every shortlisted engineer comes with both scorecards filled in. You see them before you meet anyone.

Who runs the assessments?

Senior engineers who build production software with these tools every day run both assessments. Recruiters don't score candidates.

Why not a take-home or a coding test?

Every take-home is model-written now, and LeetCode-style tests measure pattern recall under a clock. Neither shows the fundamentals or the agent skills.

Let's talk

See the scorecards before you meet anyone

Tell us the role. We send named engineers with both scorecards filled in within five working days.

Reviewed onClutch4.9 out of 5 from 36 reviews
ISO 27001
Certified

Book thirty minutes with Dale

The calendar is provided by HubSpot, which sets its own cookies. Load it here, or book on HubSpot's page.

Open booking page