How do you test software you didn't design

Practice3 min read

Testing software you didn't design means you lack the mental model the author used. With AI-generated code, there was no human author, so that model may not exist at all. The practical consequence is that coverage metrics become unreliable as proxies for confidence.

Why coverage misleads here

Line and branch coverage tell you which code ran during tests. They say nothing about whether the tests were written to probe the right behaviour. When a model generates both the implementation and the test suite, the tests tend to reflect the same assumptions that shaped the code. Gaps in the model's understanding appear in both places simultaneously, so coverage figures look healthy while meaningful scenarios go untested.

This is not a theoretical risk. AI models generate plausible-looking code. Plausible and correct are not the same thing, and a test written by the same model that wrote the function will rarely surface the difference.

What to do instead

The shift is from measuring execution to reasoning about behaviour.

Start with the specification, not the code. Write tests from requirements or expected behaviour before reading the implementation. If no specification exists, write one, even briefly, before opening the generated file. This prevents the code from anchoring your expectations.

Test the boundaries the model was unlikely to consider. Models tend to handle the happy path well. Edge cases involving empty inputs, concurrent writes, locale differences, integer overflow, or missing permissions are more likely to be thin. Enumerate these independently.

Use property-based testing where practical. Rather than asserting specific outputs, assert invariants: a sort function should always return a list of the same length; a pricing function should never return a negative value. Property-based tools generate hundreds of inputs automatically and surface assumptions the model baked in silently.

Read the code for intent, not just correctness. Before trusting the test suite, trace through the implementation and ask what the code is optimised for. AI-generated code sometimes solves a slightly different problem than the one stated, particularly when the prompt was ambiguous. Understanding the actual behaviour, rather than the intended behaviour, changes what you test.

Treat external calls with extra scepticism. Generated code that calls APIs, databases or queues often mocks those dependencies in tests. Check whether the mocks reflect how the real dependency actually behaves, including failure modes and rate limits, not just the success case.

The confidence question

Coverage is a proxy for confidence because, in code a human wrote, the person who wrote it usually also designed the tests to catch the things they were worried about. That relationship breaks down when the code and tests share a single automated origin.

Confidence in AI-generated code has to come from somewhere else: from understanding what the code is doing, from tests written against the specification rather than the implementation, and from systematic attention to the categories of failure that models reliably underweight.

This is related to the broader question of how to review AI-generated code. Code review and testing are different activities, yet when the code has no human designer, they need to compensate for each other more than usual. An engineer who can only run a linter and check coverage is not equipped for this.

It also connects to what changed about code review more generally. The reviewing engineer is now the primary source of design judgement, not just an approver of someone else's.

What we test for

Three of our scorecard criteria speak directly to this. AI output verification assesses whether an engineer catches hallucinated APIs, subtly wrong logic or security gaps before generated code ships. Judgement by risk looks at whether review effort scales with consequence, not with how confident the agent sounded. Ownership covers tests, observability and post-release debugging, regardless of who wrote the code. Both sessions test these under realistic conditions, with and without AI tools. See how we vet.

Short answers

Why is 100% coverage not enough for AI-generated code?

Coverage measures which lines ran, not whether the right questions were asked. When the same model writes the code and the tests, both share the same blind spots. High coverage can coexist with large untested categories of behaviour.

Should you write tests before or after reading AI-generated code?

Before, where possible. Reading the implementation first anchors your expectations to what the model produced. Writing tests from the specification first forces you to define correct behaviour independently, which surfaces discrepancies more reliably.

What kinds of failures do AI-generated tests most commonly miss?

Edge cases with empty or malformed inputs, concurrent access, external dependency failures, locale or timezone sensitivity, and permission boundaries. Models optimise for the happy path; systematic testing needs to cover the cases that are unlikely but consequential.

Let's talk

Get a shortlist within five working days

You share the roles and the stack in a short form or a thirty-minute call. Within five working days you get named senior engineers to review, each with both scorecards.

Reviewed onClutch4.9 out of 5 from 36 reviews
ISO 27001
Certified

Book thirty minutes with Dale

The calendar is provided by HubSpot, which sets its own cookies. Load it here, or book on HubSpot's page.

Open booking page