cavendish
TestedValidatinggate · ReliabilityNear · 1–3 years×4 sightings

Fully Agentic QA

Agents can own test design, execution, triage and regression on a real codebase, but the unit that survives is an agent-run QA function with a human owning the risk model — not a suite that writes itself and a QA team that goes away.

Experiment run, measured result. The only tier that becomes a recommendation.

Join with…

Confidence

61%human-committed

Expiry

21duntil review · 24 Sep 2026

Lead time

+4moahead of mainstream awareness

Ownership

SKSam Kowalczykfortnightly cadence

Where it is

The pilot on a regression-heavy internal repo (x-agentic-qa-pilot) settled the easy half: agents write and maintain test suites faster than people, and they triage flaky failures better than any of our engineers because they do not get bored. The hard half is unsettled. Agent-designed suites over-index on what the code does and under-index on what the business would be embarrassed by, and the coverage number climbs while the escape rate does not fall. Every vendor now sells 'autonomous QA'; nobody publishes escape rates. The field is validating because we have one measured result and it points both ways.

Why a Quantium decision hinges on it

Quantium's delivery model bills QA as a distinct line on most engineering engagements, and three clients have asked whether it should still be there. The answer determines pricing, staffing and what the firm says about its own agentic delivery claims. It also sets the reliability bar for Dark Harness: unattended production is not credible until QA is unattended first.

Field attributes

StateValidating
GateReliability · possible, not yet dependable enough
OriginPressure
Measurablepartial
Audience · TLPpractice
Horizonnear
Opened18 Mar 2026
Mainstream20 Jul 2026
Last validated27 Aug 2026
Sightings4

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01On a 14k-test internal repo, an agent-run QA loop cut suite maintenance time by 71% and triaged 92% of flaky failures to the correct cause without a human (x-agentic-qa-pilot).
  • 02Agent-authored regression tests caught 4 of 5 seeded defects that the human-written suite missed; they missed 2 of 5 that the human suite caught. Complementary, not substitutive.
  • 03Escape rate to staging did not move over eight weeks (0.9 → 0.8 per hundred merged PRs), despite line coverage rising from 64% to 88%.
What is hype
  • 01'Coverage went from 60% to 95%.' Coverage is the number agents move most easily and the one least correlated with escapes in our data.
  • 02Vendor 'autonomous QA' demos on greenfield repos. Nobody demos on a fifteen-year-old Java monolith with a hand-maintained fixture database.
  • 03'QA is dead.' The role changes; the accountability does not, and someone still signs the release.
What would have to be true
  • 01An agent that can derive a risk model from business context — what would embarrass the client — and weight test design by it. We have not seen one.
  • 02Escape rate falling, measured on at least two client codebases, not coverage rising.
  • 03A regulator or auditor accepting agent-authored test evidence for a CPS 230 control without a human attestation on each suite.
What we would do
  • 01Re-run the pilot on a client codebase under a Type 3 with escape rate as the only primary metric.
  • 02Publish a not-yet recommendation: agents own test maintenance and triage; humans own test design for risk-weighted paths.
  • 03Stand up an escape-rate ledger in Nightingale so the firm's own delivery QA becomes measurable before we advise clients on theirs.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
−71%
maintenance effort
Finding·band 1Tested

Agentic QA pilot: 71% less maintenance, escape rate unchanged

Eight weeks of agent-run QA on a 14k-test internal repo with seeded defects. Maintenance and triage effort fell sharply; coverage rose to 88%; escapes to staging stayed flat. Agent and human suites caught different defect classes.

extracted claimAgentic QA cuts maintenance effort by more than half without moving the escape rate.
Lab · x-agentic-qa-pilot · Sam Kowalczyk27 Aug 2026
detector · bleeding edge
−38%
test-eng openings
Job posting·band 2Signal

Big-four bank: 'test engineer' openings down, 'quality engineer, agent systems' up

Argus scan of AU banking engineering postings over six months. Test-engineer titles fell 38%; a new title carrying 'agent' and 'quality' appeared at three of four majors. Inference: the role is being re-titled, not cut.

Job boards · AU banking1 Aug 2026
detector · early adoption
Post·band 2Signal

'The agent tested everything except the thing that would have got me fired'

Widely shared post-mortem: an agent-maintained suite at 91% coverage missed a refund-ordering bug because nothing in the repo said refunds were the risky path. Argues the risk model has to be an input, not an inference.

Personal blog · A staff engineer at a payments company25 Jul 2026
SK?dropped 4
Talk·band 3Signal

Panel: 'What does a QA team do in 2027?'

Demand-band signal. Four of five panellists described a smaller team owning risk models and release attestation, with agents underneath. The fifth said the team goes away. Audience vote 70/30 for the first view.

YOW! Sydney20 Jul 2026
LFdropped
Release·band 1Signal

Three test-tooling vendors ship 'autonomous QA agent' in the same fortnight

All three claim end-to-end ownership from test design to sign-off. None publishes escape rates; two demo on greenfield Next.js apps. Naming event: 'autonomous QA' replaced 'AI-assisted testing' across all three.

Vendor changelogs2 Jul 2026
detector · bleeding edge 2
40%
claimed cut
Analyst·band 3Signal

Analyst note: 'Autonomous testing will remove 40% of QA headcount by 2028'

Headline number with no methodology beyond vendor interviews. Kept as the strongest overstatement in circulation; three clients cited it in the same month.

Industry analyst briefing30 Jun 2026
RMdropped 3
31
projects
Paper·band 1Signal

Coverage Is Not Confidence: LLM-Generated Test Suites and Defect Escape

Evaluates LLM-generated suites on 31 open-source projects with historical defect data. Coverage gains of 20–35 points correspond to escape-rate reductions statistically indistinguishable from zero. Mutation score tracks escapes; coverage does not.

extracted claimLLM-generated coverage gains do not predict defect-escape reductions; mutation score does.
arxiv.org · Nakamura, Oyelaran et al.19 Jun 2026
MTSKdropped 3
Finding·band 1Tried

Logged from Claude Code: agent triaged a two-year-old flaky test to a timezone fixture

Product engineer logged from a session: a test that had been quarantined since 2024 was traced by the agent to a fixture with a hard-coded AEDT offset. Fixed in one PR. Tried tier; one instance.

MCP · log_finding · Ollie Grant11 Jun 2026
OGdropped
Client question·band 3Signal

'If your agents write the tests, why is QA still a line on the SOW?'

Asked in a commercial review by a telco client's procurement lead. Logged by the delivery lead; the question that moved the field from candidate to elected under pressure.

Engel · telco engagement22 May 2026
MLdropped 3
41%
defect overlap
Benchmark·band 2Tried

SWE-Test-Bench: seeded-defect detection by generated suites

Public harness that seeds defects into real repos and scores generated suites on detection. We ran the pilot repo through it; the agent suite and human suite had a 41% overlap in detected defects.

github.com6 May 2026
MTdropped 2
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 5 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Agent-run test maintenance and flaky-triage cut QA effort by more than half on a regression-heavy codebase with no rise in escapes.

Testedc-fully-agentic-qa-1dalton-0.427 Aug 2026Lab · x-agentic-qa-pilot, MCP · log_finding
81%

Line coverage produced by agents is uncorrelated with escape rate; coverage is the wrong headline metric for agentic QA.

Testedc-fully-agentic-qa-2dalton-0.427 Aug 2026Lab · x-agentic-qa-pilot, arxiv.org
77%

Agent-written and human-written suites catch different defect classes; the union beats either alone by a margin worth paying for.

Testedc-fully-agentic-qa-6dalton-0.427 Aug 2026Lab · x-agentic-qa-pilot, github.com
72%

Agent-designed suites systematically under-test business-risk paths because the risk model is not in the repo.

Testedc-fully-agentic-qa-3dalton-0.427 Aug 2026Lab · x-agentic-qa-pilot, Personal blog
69%

The QA role is being re-titled, not removed: openings for 'test engineer' are falling while 'quality engineer, agent systems' openings rise.

Assessedc-fully-agentic-qa-5dalton-0.45 Aug 2026Job boards · AU banking, YOW! Sydney
58%

A fully unattended QA function — design through sign-off — is deployable on enterprise codebases within 18 months.

Assessedc-fully-agentic-qa-4dalton-0.414 Jul 2026Vendor changelogs, Industry analyst briefing
24%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 50% → 61%.

runs compare claim sets, never prose
What we said · run 4

Pilot concluded. Effort saving is real, escape rate is flat, coverage is a vanity metric. The role is re-shaped around the risk model, not removed. Full-replacement thesis sent to the graveyard.

61%
Changed since run 3
  • Agent-run test maintenance and flaky-triage cut QA effort by more than half on a regression-heavy codebase with no rise in escapes.
  • Line coverage produced by agents is uncorrelated with escape rate; coverage is the wrong headline metric for agentic QA.
  • Agent-designed suites systematically under-test business-risk paths because the risk model is not in the repo.
  • Agent-written and human-written suites catch different defect classes; the union beats either alone by a margin worth paying for.
  • c-fully-agentic-qa-4 ↓ 0.38 → 0.24
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · AW
high

Changes the staffing model on every engineering engagement the firm runs.

Timeline

committed · SK
0–18mo

Maintenance and triage are usable now; design and sign-off are not.

TAM

agent-estimated
$1B–10B

Agent-estimated from global QA outsourcing spend and test-tooling revenue. Uncommitted.

Cost

committed · SK
low

One engineer on the pilot repo for six weeks; tokens under $2k.

Cost of being wrong

committed · LF
high

If we tell a bank QA is unattended and it is not, the escape lands in production under CPS 230.

Demand

committed · ML
high

Three engagements asked whether QA should remain a billed line; one RFP scores 'AI-driven testing' explicitly.

Workforce readiness

agent-estimated
medium

Delivery pods can run maintenance/triage today; nobody outside the lab has written a risk-weighted test plan for an agent. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Banking
relevant

CPS 230 makes test evidence a control artefact. Agent-authored evidence is only useful if an auditor will accept it, which none has yet.

Mechanism · Agents maintain and run suites; a named engineer attests each release against a risk-weighted test plan.

CD committed by Claire Duboiscommitted
Telco
relevant

Provisioning and billing regression suites are large, brittle and mostly maintenance. This is where the 71% lands first.

Mechanism · Agent-run maintenance and flaky triage over the existing suite before any redesign.

ML committed by Marcus Leecommitted
Government
watch

Procurement panels still specify tester headcount. Until the buying template changes, the saving is not bankable.

Mechanism · Would need DTA guidance on agent-produced assurance evidence.

Agent draft · awaiting a sector owneragent-estimated
Retail & FMCG
relevant

Promotional-pricing logic breaks in ways the code never predicts. Agents wrote 400 tests and none exercised a stacked-discount path.

Mechanism · Human-authored risk model feeds the agent's test design; agents own everything below it.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against: we measured the easy metrics on a codebase we already understood. Maintenance and triage were always the cheap part of QA; the expensive part is knowing what to test, and the pilot showed agents do not know that. A 71% saving on 30% of the work is a 21% saving being sold as a transformation.

  • Escape rate did not move. On the only metric that matters to a client, eight weeks of agentic QA produced no measurable change.
  • The pilot repo has a complete, honest test suite already. Most client codebases do not, and an agent maintaining a bad suite maintains it faster.
  • The 'union beats either' result depends on keeping the human suite, which means keeping the humans. The economics disappear if the headline is staffing.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis weakened

Source diversity

  • Software engineering research30%
  • Practitioner posts25%
  • Vendor / analyst20%
  • Internal / Engel25%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsSKMTOGMLLF

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Can a risk model be expressed in a form an agent can weight test design by, and who in a delivery pod writes it?
  2. 02Does the escape-rate result hold on a client codebase without a pre-existing honest suite?
  3. 03Will an APRA-regulated entity's auditor accept agent-authored test evidence, and under what attestation?