What does inference actually cost right now?
Blended across our six instrumented patterns, 1 September: $0.9–1.6 per thousand agent turns on the open-weight tier, $4–11 on frontier mid-tier. Auto-refreshed from the ledger; not yet human re-committed.
Tier is not strength
Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.
Machine-readable target
{
"configKey": "ledger.cost.blended_per_kturn"
}Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.
Body
As of 1 September 2026, from the token cost ledger (auto-refreshed; Priya has not re-committed this cycle): list prices sit at $0.40–0.90 per million input tokens for the open-weight 30B-class tier via AU-region hosts, roughly $3 per million input for frontier mid-tier and $15 for frontier top-tier, with output tokens at 4–5× input across all three. What matters more is the blended cost per agent turn on real patterns: $0.9–1.6 per thousand turns on open-weight, $4–11 on frontier mid-tier, with the range driven by cache hit rate and output length.
Evidence: x-token-cost-ledger, running continuously since June across six client patterns (c-token-cost-reduction-1). The blended figure has fallen 23% since June, split roughly evenly between list-price cuts and routing more work to the open-weight tier (c-token-cost-reduction-3). Caching contributes 30–45% on input, not the 60% headline; see r-prompt-caching.
Caveat: this answer is refreshed from the ledger every fortnight and carries a 21-day half-life. The auto-refresh updates the numbers; it does not re-check whether the pattern mix still represents the firm. Two providers changed cache-write pricing in the last quarter and one may again. Quote a band, never a point, to a client.
What it rests on
Prompt caching yields 30–45% per-task savings on stable prefixes over 2k tokens and near nothing below; it is not a universal 60%.
Batching halves cost on any workload that tolerates hours of latency, with no engineering beyond a queue.
Field
Cost redux on tokensExperiment · validated
Token cost ledger across six client patternsGraveyard · refuted
Prompt caching as a universal 60% cost cut “Sixty percent, on the vendor's workload.”