Somewhere in the last year, one question started opening almost every call we have with engineering leaders: are we behind? Is everyone else already running parallel agents? Has anyone actually dropped the take-home, or do they just post about it? The honest answer, until now, has been that everybody is guessing, including the people posting most confidently about it.
The Engineering Leader Benchmark opened this week.
The mechanics are simple. Twenty questions on how your team works with AI, how you hire, how you verify and how you size. The point is seeing your answers next to those of the other engineering leaders who filled it in; the score itself matters less. You'll find out whether the thing you assumed everyone struggles with is actually just you, and where other teams are further along. We publish the results aggregated and anonymised, and everyone who contributes gets the report and can see where they sit within the benchmark compared to their peers.
The four areas aren't arbitrary. They're the four places where, in two years of building with these tools and assessing the engineers who use them, we've watched the way of working lag furthest behind the tooling.
How your team works with AI is the productivity area. The whitepaper behind this benchmark opens with three engineers on the same backlog: one who doesn't use AI, one prompting in a single window, one moving between four git worktrees with an agent in each. By lunch the third has done what the first will finish on Thursday. The questions here establish which of those three your team actually resembles, as opposed to which one it identifies with.
How you hire is the filter area. Every take-home submitted anywhere is model-written now, so the artefact carries no information about the person, and pattern-recall tests measure precisely the work agents absorbed. The questions here get at whether your process ever watches a candidate work, with and without the tools, or whether it still evaluates output.
How you verify is the trust area. When most shipped code is generated, the engineer's real contribution is validation: catching the hallucinated API, the subtly wrong logic, the security gap, before it ships rather than after. Teams rarely know how much of what they merge has actually been read. The questions here make that measurable, and this is the area where we expect the most uncomfortable answers.
How you size is the planning area. On our builds, a project that needed ten engineers in 2024 runs on four today. Whether your own ratio has moved, and whether your 2027 headcount plan reflects it, says more about your operating model than any tooling survey.
Twenty questions, four areas, a few honest minutes. What you get back is the one thing nobody in this market has had: your position relative to peers answering the same questions about the same shift, rather than another opinion about what teams "should" be doing.
The first report is being compiled from early responses now, and contributors read it before anyone else. If you run an engineering team and you're locking 2027 plans this quarter, this is the cheapest reality check available to you.
Apply for the benchmark report. The full argument behind the questions, including both assessments we run on engineers, is in The Next 10X Engineer.



