cavendish
Standing answerAssessedstrength · moderate

Is the vendor's 'agentic' claim real?

Usually not in the sense the deck implies. Ask four questions: does it plan across more than one tool call, can it recover from a failed call, what happens on ambiguity, and where is the eval. As of 14 August, most vendors fail two of four.

Tier is not strength

Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.

Evidence tierAssessed
Strengthmoderate
OwnerLFLena Fischer
Last validated14 Aug 2026
Review by13 Sep 2026
Half-life30 days
Citations19
Asked58× this quarter
VerticalsCross-sector, Banking, Government
decay10d until review

Machine-readable target

{
  "taskType": "vendor-assessment",
  "configKey": "assurance.vendor.agentic_checklist"
}

Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.

Body

As of 14 August 2026: treat 'agentic' as unverified until the vendor answers four questions on a live demo, not a video. Does the product choose and sequence more than one tool call without a fixed script? Does it recover from a failed tool call rather than halt or hallucinate? What does it do when the task is ambiguous — ask, guess, or stop? Where is the eval, and who calibrated it? Of fourteen vendor claims assessed this year, nine failed at least two of the four (c-table-stakes-2).

Evidence: the table-stakes validation runs (two so far) and the eval harness field's assessment method. The most common pattern is a fixed workflow with a language model at one step, relabelled. That is fine and often what the client needs; it is not an agent, and it should not be priced or risk-assessed as one (c-table-stakes-3).

Caveat: the bar moves. Two of the nine that failed in March passed in August. Ask again rather than relying on the previous answer, and log the answer so the next person does not have to.

LFSigned Lena Fischer · Red team & assurance · 14 Aug 2026

What it rests on

Evals are claimed by half the competitor set and demonstrated publicly by none; they are six to twelve months from table stakes.

Assessed c-table-stakes-2
66%

'Proprietary orchestration' claims in competitor pitches map to open-source harnesses in the majority of cases examined.

Tried c-table-stakes-3
60%

Field

Deciding Table Stakes