Graveyard · killed by Cavendish
What we killed.
With names and causes.
Refuted, abandoned, superseded, rejected — each with a bylined author, a plainly stated cause, and a flag if the constraint that killed it might move. Abandoned is never dressed as refuted. Every entry keeps its full lineage.
- Killed
- 0
- Resurrectable
- 0
- Returned to pile
- 0
Mean lifespan 100 days. Corrections counted on the scorecard, without a target.
- SupersededTested
Vector-store memory as the agent's long-term memory
“Remembered everything, retrieved nothing.”
10 Feb 2026 — 21 Aug 20266 monthsCause · The memory bake-off answered the question with a different design: structured episodic store with summarised recall beat raw vector recall on precision and latency above 50k tokens of history.
TOTom OkaforAgentic Memory System - AbandonedTried
Realtime voice agents replace tier-1 IVR
“Hung up before the answer.”
7 Apr 2026 — 15 Jul 20263 monthsCause · The telco partner withdrew the test trunk on day 24 of a planned 40-call measurement and the pair was reassigned; 11 calls were measured, which is not enough to answer the hypothesis either way.
PRPriya Raman resurrectableVoice and Vision - RefutedTried
Self-reported productivity surveys as ROI evidence
“Everyone said it was working.”
14 Oct 2025 — 14 Jul 20269 monthsCause · In the first arm of the Nightingale run, self-reported time saved per developer correlated at 0.11 with measured cycle-time change over the same ten weeks, and over-stated it by a median factor of 3.4.
JPJun ParkROI - RejectedSignal
World-model simulation for retail demand planning
“Simulated a world that had no shelves in it.”
16 Jun 2026 — 10 Jul 202624 daysCause · Not selected: no published world model operates on tabular demand series, the proposed run would have been a conventional forecasting comparison with a new name, and the field is 'next' with a stated breakthrough gate that has not moved.
AWAdam Witanowski resurrectableWorld Models - SupersededTested
One commercial gateway over every model garden
“Standardised on someone else's roadmap.”
2 Dec 2025 — 30 Jun 20267 monthsCause · The routing experiment showed the value sits in the routing policy and the cost ledger, both of which we could not get out of the vendor gateway, so the recommendation became a gateway we control.
PRPriya RamanAI Gateway - AbandonedTried
Ambient Slack agent that acts unprompted
“Waited for permission. Still waiting.”
12 May 2026 — 26 Jun 20261 monthCause · The security review needed for a bot with read access to firm channels was not completed inside the run window, so the agent never went live and nothing was observed.
TOTom Okafor resurrectableAmbient Agents - RefutedTried
Template guardrails stop org slop
“Guarded the template. Not the fork.”
24 Feb 2026 — 5 Jun 20263 monthsCause · After three months with mandated project templates, the count of unowned AI-built internal tools rose from 41 to 67; builders forked the template once and never took the updates.
SKSam KowalczykCitizen Developers and Org Slop - SupersededAssessed
On-prem H100 cluster pays back inside 18 months
“Payback period outran the price list.”
18 Nov 2025 — 22 May 20266 monthsCause · Two AU-region hosted price cuts in Q1 and Q2 moved the breakeven to over four years before we finished the model; the standing answer now carries the calculation and its refresh date.
PRPriya RamanOn-Prem Inference - AbandonedTried
Plain OAuth scopes are enough for agent delegation
“Never got a tenant to fail in.”
24 Mar 2026 — 19 May 20262 monthsCause · No test tenant was made available in the run window, so the claim that native identity-provider scopes can express per-task agent delegation was never exercised against a real directory.
LFLena Fischer resurrectableAuth Broker - RefutedTested
LLM-as-judge without human calibration
“Agreed with itself, mostly.”
13 Jan 2026 — 8 May 20264 monthsCause · Against a five-person human panel on 600 paired items, an uncalibrated frontier judge reached Cohen's κ of 0.41 to 0.58 across three task families, below the 0.8 the hypothesis required and below the 0.6 kill line on two of them.
MTMei TanakaEval Harnesses - RejectedSignal
Agent-written test suites replace QA engineers
“Could not be wrong, so could not run.”
14 Apr 2026 — 28 Apr 202614 daysCause · Not selected: the hypothesis as written is a headcount claim with no measurable kill condition, and the narrower pilot that could be preregistered (x-agentic-qa-pilot) was selected in its place.
AWAdam WitanowskiFully Agentic QA - RefutedTried
Unattended agent PRs merged on green CI
“Green CI, red incident.”
10 Mar 2026 — 3 Apr 202624 daysCause · In a three-week trial on the lab's own tooling repo, 4 of 31 unattended merges passed CI and broke behaviour the suite did not cover, including one that silently changed a cost calculation.
SKSam Kowalczyk resurrectableAI-SDLC - RefutedTested
Prompt caching as a universal 60% cost cut
“Sixty percent, on the vendor's workload.”
4 Nov 2025 — 27 Mar 20265 monthsCause · Across six client patterns the measured saving from prompt caching ranged from 4% to 44%, and the 60% figure only appeared on workloads with a stable prefix over 2k tokens and a call rate high enough to keep the cache warm.
PRPriya RamanCost redux on tokens - RejectedSignal
Sell a quantum-readiness audit this year
“Right question, wrong decade, wrong department.”
3 Mar 2026 — 17 Mar 202614 daysCause · Killed on inspection: no client has asked, the ASD's published migration timeline gives a 2030 horizon for most in-scope systems, and the substance of the audit is a crypto inventory that the firm's security practice already sells.
AWAdam WitanowskiQuantum Encryption - AbandonedTried
Multi-agent debate improves reasoning on our tasks
“Adjourned, not decided.”
20 Jan 2026 — 31 Jan 202611 daysCause · A Type 2 run stopped on day 11 when the owner was pulled onto the memory bench and the claims-triage eval it was using was re-baselined under a new extractor, leaving the two completed conditions incomparable with the planned third.
TOTom OkaforSelf-Organising Agents - RefutedTried
Fine-tuned 7B beats frontier on banking classification
“Beat the baseline it was allowed to pick.”
21 Oct 2025 — 16 Jan 20263 monthsCause · On a like-for-like eval with a properly prompted frontier baseline, the fine-tuned 7B trailed by 6 F1 points; the earlier 'win' had been against a frontier model with a one-line prompt.
MTMei Tanaka resurrectableSLM / Edge / Tuning