cavendish
Auto-refreshed on schedule. No human has re-committed this refresh yet. Treat accordingly.
Standing answerTestedstrength · moderate

What does inference actually cost right now?

Blended across our six instrumented patterns, 1 September: $0.9–1.6 per thousand agent turns on the open-weight tier, $4–11 on frontier mid-tier. Auto-refreshed from the ledger; not yet human re-committed.

Tier is not strength

Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.

Evidence tierTested
Strengthmoderate
OwnerPRPriya Raman
Last validated1 Sep 2026
Review by22 Sep 2026
Half-life21 days
Citations48
Asked184× this quarter
VerticalsCross-sector
decay19d until review

Machine-readable target

{
  "configKey": "ledger.cost.blended_per_kturn"
}

Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.

Body

As of 1 September 2026, from the token cost ledger (auto-refreshed; Priya has not re-committed this cycle): list prices sit at $0.40–0.90 per million input tokens for the open-weight 30B-class tier via AU-region hosts, roughly $3 per million input for frontier mid-tier and $15 for frontier top-tier, with output tokens at 4–5× input across all three. What matters more is the blended cost per agent turn on real patterns: $0.9–1.6 per thousand turns on open-weight, $4–11 on frontier mid-tier, with the range driven by cache hit rate and output length.

Evidence: x-token-cost-ledger, running continuously since June across six client patterns (c-token-cost-reduction-1). The blended figure has fallen 23% since June, split roughly evenly between list-price cuts and routing more work to the open-weight tier (c-token-cost-reduction-3). Caching contributes 30–45% on input, not the 60% headline; see r-prompt-caching.

Caveat: this answer is refreshed from the ledger every fortnight and carries a 21-day half-life. The auto-refresh updates the numbers; it does not re-check whether the pattern mix still represents the firm. Two providers changed cache-write pricing in the last quarter and one may again. Quote a band, never a point, to a client.

PRSigned Priya Raman · Research engineer · inference · 1 Sep 2026

What it rests on

Prompt caching yields 30–45% per-task savings on stable prefixes over 2k tokens and near nothing below; it is not a universal 60%.

Tested c-token-cost-reduction-1
86%

Batching halves cost on any workload that tolerates hours of latency, with no engineering beyond a queue.

Tested c-token-cost-reduction-3
83%

Field

Cost redux on tokens

Experiment · validated

Token cost ledger across six client patterns

Graveyard · refuted

Prompt caching as a universal 60% cost cut Sixty percent, on the vendor's workload.