Deciding Table Stakes
right to play, as opposed to moats — i.e. skills
By 2026 skills, harnesses, evals and agent delivery are the price of being in the room — every competitor claims them and every client assumes them — so none of them is a moat; the only defensible position is measured results on the client's own data and the credibility to say no with evidence.
A validation run. Researched position, no experiment.
Confidence
56%human-committedExpiry
27duntil review · 30 Sep 2026Lead time
—not yet mainstream · opened 2 Jun 2026Ownership
AWAdam Witanowskimonthly cadenceWhere it is
Opened on conviction by the lab director in June after an RFP debrief in which an insurer said every firm they had spoken to 'does agents'. Argus positioning deltas across Accenture, Deloitte, QuantumBlack and Palantir since then show 'agentic' on all four AU practice pages, 'evals' on two, and a measured client outcome on none. A competitor's 'proprietary orchestration' pitch mapped to an open-source harness when Dalton read it; the director rebuilt a public 'agent factory' demo in three hours with stock tooling. A foundation lab's open skills spec has made skill libraries a commodity in one release. The skills gate is the honest one: the firm has to have all of this, cannot differentiate on any of it, and has to build it while pretending it is special. The disconfirming case — that consulting moats were always relationships and distribution and capability was never the point — is carried at moderate confidence and is the reason the field has an exec sponsor.
Why a Quantium decision hinges on it
This is the field that decides what the lab is for. If agents, evals and harnesses are table stakes, the lab's job is to get the firm to the table fast and cheaply, and the differentiated work is Nightingale measurement and the graveyard — being the firm that can show what worked and what did not on real data. If they are moats, the lab should be building proprietary tooling. The RFP scoring evidence says the first; the sponsor's conviction says the second is what wins pitches this year. The answer sets next year's budget split between capability and measurement.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Argus positioning delta, June 2026: 'agentic' present on 4 of 4 competitor AU practice pages; 'evaluation' on 2 of 4; a named, measured client outcome on 0 of 4.
- 02A competitor's 'proprietary agent orchestration' pitch language mapped by Dalton to an open-source harness with a naming layer; the director rebuilt the public demo in three hours with stock tooling (tried tier).
- 03A foundation lab published an open skills specification and marketplace in April; three competitors' 'skill libraries' now implement it.
- 04Competitor hiring for 'AI evals engineer' roles rose from near zero to 23 open AU roles in six months — evals moved from differentiator to job description.
- 01'Proprietary agent platform' claims from consultancies. Every one we have examined is a harness plus a brand.
- 02'Thousands of agents deployed.' Counts of things built, not things running, and never a measured outcome.
- 03The lab's own temptation: that a good skills library is a moat. It was, for about a quarter.
- 01RFP scoring that rewards measured outcomes over capability claims. Currently procurement panels score the claims, and a measured 'no' can lose to an unmeasured 'yes'.
- 02A Nightingale phase-2 result on the firm's own book — lead time, win rate on measured pitches — that shows measurement wins work. That is the evidence the sponsor's conviction is waiting for.
- 03At least one competitor conceding the capability layer is a commodity in public. When the first one says it, the market moves.
- 01Publish pos-table-stakes as the firm's position and keep sa-vendor-claim-agents current for every sales conversation that starts with 'the vendor says'.
- 02Fund capability building to parity, not beyond: skills, harness, evals, agent delivery, at the cheapest credible level. Spend the difference on measurement.
- 03Run a Type 2 each quarter: rebuild the most-cited competitor demo with stock tooling and log the hours. The gap is the moat, measured.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Argus positioning delta: 'agentic' on 4 of 4 competitor AU practice pages; measured outcome on 0
Computed from the claim layer across Accenture, Deloitte, QuantumBlack and Palantir AU practice pages, quarterly since 2025-09. 'Agentic' went from 1 of 4 to 4 of 4 in three quarters; 'evaluation' from 0 to 2; a named client with a measured outcome stayed at 0 throughout. The delta that opened the field.
extracted claimAgentic delivery claims have reached saturation across the AU competitor set inside three quarters, and none is backed by a published measured outcome.

'Consulting moats were never capabilities'
Argues that every consulting moat in history has been trust, distribution and incumbency; that capability claims are theatre both sides perform; and that a firm owned by a retailer has the only moat that matters in retail. Strongest disconfirming voice and the red team's spine.

'The model is not the moat. Neither is the harness.' — competitor keynote
A QuantumBlack partner says on stage that the capability layer is a commodity and the differentiation is 'outcomes on your data'. The first competitor concession we have recorded; it is also exactly our position, which either confirms it or makes it table stakes too.

'AI Evaluation Engineer' — 23 open roles across AU competitor firms
Argus hiring scan. Near zero in January; 23 in July across four firms. Inference: evals are being staffed as a delivery function, which puts them on the table-stakes side within a year.

Logged from Claude Code: rebuilt a competitor's public 'agent factory' demo in three hours with a stock harness and open skills
The director reproduced the flow shown in a competitor's public demo video — intake, triage, drafting, handoff — using a stock harness, the open skills spec and no bespoke code. Three hours including the video. One run, tried tier, and a demo is not a programme.

Competitor pitch excerpt: 'proprietary agent orchestration layer' — mapped by Dalton to an open-source harness
The sponsor dropped a de-identified excerpt from a competitor's pitch to a shared client. Dalton matched the described features one-for-one to an open-source harness's documentation. Aggregate only; no engagement named.

'Every firm we spoke to says they do agents. What do you do that they don't?'
Asked by a COO at a debrief the firm lost. The sponsor logged it verbatim. The question the field exists to answer, and the reason it has an exec sponsor.

'Skills are the new prompts — and just as easy to copy'
Argues that skill libraries, like prompt libraries before them, have a shelf life of one open spec. Written three weeks after the open skills release; correct so far.

Foundation lab publishes an open skills specification and marketplace
Portable skill packaging with a public directory. Within two months three competitors' 'skill libraries' were implementations of it. The release that turned the lab's own skills library from an asset into a table stake.

Accenture announces multi-thousand-headcount 'agentic services' practice and an agent factory
Headcount, a factory metaphor and a partnership list. No outcome figures. Argus records the move from 'AI' to 'agentic' in the firm's naming; the largest competitor now uses the same word as everyone else.
extracted claimThe largest consultancy has rebranded its AI practice around agents at scale.
Claims · 4 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
Every top-tier competitor in Australia now publicly claims agentic delivery; the claim no longer discriminates in an RFP.
Evals are claimed by half the competitor set and demonstrated publicly by none; they are six to twelve months from table stakes.
'Proprietary orchestration' claims in competitor pitches map to open-source harnesses in the majority of cases examined.
The only capability a client cannot obtain from any competitor is a measured outcome on their own data, published with what did not work.
Capability is irrelevant to consulting moats; relationships and distribution decide, and the firm's moat is the Woolworths relationship.
Position history · the diff is the product
3 validation runs against a fixed brief. Confidence 45% → 56%.
Position drafted: parity on capability, differentiation on measurement and the graveyard. Relationship counter-case rated higher than expected; verdict in doubt until Engel can show measured pitches win. Position and standing answer published with the doubt stated.
- The only capability a client cannot obtain from any competitor is a measured outcome on their own data, published with what did not work.
- Capability is irrelevant to consulting moats; relationships and distribution decide, and the firm's moat is the Woolworths relationship.
- c-table-stakes-3 ↑ 0.52 → 0.60
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · RMSets the capability-versus-measurement budget split for next year.
Timeline
committed · AWThe skills gate closes on the firm's own build pace, not the market's.
Cost of being wrong
committed · AWWrong one way funds a proprietary platform nobody buys; wrong the other way loses pitches to firms that claim more.
Demand
committed · RMEvery pitch debrief this half has had a 'what do you do that they don't' moment.
TAM
agent-estimatedAU AI-services market; not a meaningful axis for a positioning field. Agent-estimated, uncommitted.
Workforce readiness
agent-estimatedThe firm has the skills at lab depth, not delivery depth. That is the gate. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
The positioning question is the same in every sector; the RFP language differs only in which vendor is being quoted.
Mechanism · Position paper plus the standing answer on vendor agentic claims, used in every pitch.
Bank procurement panels score capability claims on a grid; a measured 'not yet' scores below an unmeasured 'yes'.
Mechanism · Lead with the graveyard and the measured result; make the panel score evidence.
Panel arrangements reward breadth of claimed capability; the measured-outcome position may not score until the buying guidance changes.
Mechanism · Track whether the digital sourcing framework adds evidence-of-outcome criteria.
The Woolworths relationship is the disconfirming case's strongest example, so it is where the thesis gets tested hardest.
Mechanism · Compare win rate on measured pitches against relationship pitches once Engel data allows. Agent draft.
Red team · the strongest case against
The strongest case against: this is a lab justifying itself. 'Capability is table stakes, measurement is the moat' is exactly what a measurement-led lab would conclude, and the evidence is four practice pages and one rebuilt demo. Consulting has never been won on capability. It is won on who the CEO trusts, which is why a firm owned by the country's largest retailer has a moat no eval will move. The field's origin is conviction and its sponsor is the person whose conviction it is.
- —Four competitor practice pages are marketing, not capability. Their claims being undifferentiated says nothing about whether their delivery is.
- —One RFP debrief and one rebuilt demo. Three hours to rebuild a demo is not three hours to deliver a programme.
- —Measured outcomes are a moat only if clients pay for them. Engel does not yet show that a measured pitch wins more often than a confident one.
- —The field's own conclusion allocates budget to the lab that wrote it. That is a conflict, and it is not addressed by naming it.
Source diversity
- Competitors / Argus40%
- Foundation labs15%
- Practitioner and commentary20%
- Internal / Engel25%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Evals are the table stake most likely to be a brief moat; the field's own timing depends on Turing.
A measured return on client data is the one claim a practice page cannot copy.
The graveyard published is the evidence that the firm says no; it is the moat's public face.
Agent delivery at parity is the capability the firm has to reach before the positioning question is even live.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- RMRohan Mehta · Exec sponsor2 drops
- SKSam Kowalczyk · Research engineer · SDLC1 drop
- AWAdam Witanowski · Lab Director (acting)1 drop
- ?Anonymous · Anonymous drop1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Does a measured pitch win more often than a confident one? Engel can answer this and has not yet.
- 02How long is the window between a capability being claimed by one competitor and by all of them — and is it shrinking?
- 03If measurement is the moat, what stops a competitor with a bigger client base measuring more than we can?