Non-Weight-Bound Continuous Learning
Deployed agents will get better with use through retrieval-updated skills, memory and test-time adaptation long before weights learn in production; 'continual learning' as the labs mean it lands after 2029, and most of its practical value arrives earlier through the non-weight route.
A validation run. Researched position, no experiment.
Confidence
58%human-committedExpiry
22duntil review · 25 Sep 2026Lead time
—not yet mainstream · opened 9 Mar 2026Ownership
MTMei Tanakamonthly cadenceWhere it is
The lab elected this in March as a hypothesis, not a signal: clients kept asking why the agent does not improve, and the honest answer was that nothing in production learns. Since then the non-weight route has firmed up. Retrieval-updated skill libraries, episodic memory and test-time training on the current task each show measurable improvement over deployment weeks in our own experiment and in two published results. True weight-level continual learning — the model updating in production without forgetting — remains a research problem with one credible lab result and no deployment. The position is that the practical field is the non-weight one, and the weight-bound one is a Distant-horizon watch item we carry here so the two are not confused.
Why a Quantium decision hinges on it
Every multi-session agent pattern the firm delivers is judged by the client on whether it gets better. Today it does not, and delivery teams are hand-building improvement loops per engagement. A defensible default for how an agent learns from use — and an honest statement of what it cannot learn — is the difference between an agent programme that compounds and one that is re-prompted every quarter. It also determines whether the Agentic Memory and Personal Wiki investments are the destination or a stopgap.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Retrieval-updated skill library: a claims agent's first-pass accuracy rose 11 points over six weeks of deployment with no weight update (x-continuous-learning, in measurement).
- 02Test-time training on the current task's context improves long-document extraction by 4–7 points in two published results and in one lab reproduction.
- 03Weight-level continual learning without forgetting: one credible lab result on a narrow benchmark; no deployment anywhere we can find.
- 01'The model learns from your data' in vendor copy that means RAG.
- 02Fine-tuning nightly as 'continuous learning'. It is batch retraining with a forgetting problem and no eval gate.
- 03Lab announcements of continual learning that measure retention on benchmarks, not improvement on a deployed task.
- 01The skill-library gain holds past twelve weeks without the library degrading into contradictions — the run has six weeks of data.
- 02A forgetting and audit mechanism: a regulated client must be able to say what the agent learned, from whom, and remove it.
- 03For the weight route: a published deployment where a model updates in production and a held-out eval does not regress over a quarter.
- 01Publish pos-continuous-learning as the position: non-weight learning is the practical field; weight-level is Distant.
- 02Conclude x-continuous-learning at twelve weeks; promote the skill-library pattern to a recommendation if the gain holds and the contradiction rate stays under threshold.
- 03Fold the forgetting/audit requirement into the Agentic Memory forgetting mechanism rather than building it twice.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.
x-continuous-learning at six weeks: skill library lifts claims first-pass accuracy 11 points, no weight update
Interim measurement from the preregistered experiment. Claims agent with a retrieval-updated skill library, seeded from adjudicated cases, against a frozen control. Gain widens weekly; contradiction rate in the library is rising and is the kill-condition watch. Not concluded.
extracted claimA retrieval-updated skill library improves a deployed agent's accuracy over weeks without weight updates.
Two frontier labs hiring 'Continual Learning — Production' researchers
Both postings name deployment-time learning explicitly. Argus inference: the weight route is being worked on seriously at two labs; our 2029 date is an estimate, not a fact.
Logged from Claude Code: skill file for the pricing assistant contradicted itself after nine weeks
Product engineer's log: a self-updated skill file on the pricing assistant accumulated two incompatible rules for promo handling and the assistant alternated between them. Fixed by hand. Tried tier; became the contradiction claim and the experiment's kill-condition watch.
'Continual learning is solved and your model will learn from you next year'
Extrapolates the narrow lab result to production models within a year. No mechanism for audit or forgetting discussed. Kept as the disconfirming voice on our timeline.
Continual Pretraining in Deployment Without Catastrophic Forgetting: A Narrow Result
The one credible weight-level result: a model updated on a streaming corpus for 30 days with under 1 point of regression on a held-out suite. Narrow domain, small model, no deployment. The reason the weight route is Distant and not dead.
extracted claimWeight-level continual learning without forgetting is demonstrated at small scale in a narrow domain.
'We fine-tuned nightly for six months. Here is what broke.'
Practitioner post-mortem: nightly fine-tuning on production traffic regressed a held-out suite by 9 points over a quarter before anyone noticed. The reason 'just retrain' is in the hype list.
'Why is it still making the mistake we corrected in March?'
Asked by a servicing operations lead about a pilot agent. Third occurrence of the same question across two banks and an insurer. The demand signal that elected the field.
Skill Libraries as a Substitute for Weight Updates in Long-Running Agents
Agents that write and retrieve their own procedural skills match a fine-tuned baseline on three long-horizon tasks after eight weeks of use. The paper that shaped the experiment design.
extracted claimSelf-written skill libraries match fine-tuning on long-horizon tasks after weeks of use.
Test-Time Training for Long-Document Understanding: Gains and Costs
Adapting a model to the current document at inference gives 4–7 points on extraction and QA at 2–3× cost. Reproduced by the lab on a contract corpus. A real non-weight-in-production mechanism, expensive.
Claims · 5 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
A retrieval-updated skill library improves a deployed agent's task accuracy over weeks without any weight update.
Clients experience 'the agent does not improve' as a defect, and it is the most common reason a multi-session agent pilot is not extended.
Test-time training on the current task context improves long-document extraction by 4–7 points at 2–3× inference cost.
Nightly fine-tuning as a substitute for continual learning regresses held-out evals in most deployments that try it.
Skill libraries accumulate contradictions over time; without a consolidation step the gain reverses after roughly two months.
Weight-level continual learning in production without forgetting will be available from a frontier lab within the Next horizon.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 40% → 58%.
Skill library shows an 11-point gain at six weeks; contradictions are the named risk. Weight-level continual learning stays Distant. Position drafted for publication.
- A retrieval-updated skill library improves a deployed agent's task accuracy over weeks without any weight update.
- Test-time training on the current task context improves long-document extraction by 4–7 points at 2–3× inference cost.
- Skill libraries accumulate contradictions over time; without a consolidation step the gain reverses after roughly two months.
- Nightly fine-tuning as a substitute for continual learning regresses held-out evals in most deployments that try it.
- c-continuous-learning-3 ↓ 0.36 → 0.28
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · MTDetermines whether every multi-session agent compounds or resets.
Timeline
committed · MTNon-weight route is deployable within two years; the weight route is the Distant part of the thesis.
TAM
agent-estimatedAgent-estimated as a share of the agent-platform market; not separable. Uncommitted.
Cost
committed · AWTwelve-week experiment with a pair; the audit mechanism is the expensive part.
Cost of being wrong
committed · LFAn agent that learns the wrong thing from a client's users is a conduct problem in banking; the audit mechanism is the control.
Demand
committed · CDAsked in two banking and one insurance engagement as 'why doesn't it get better'; nobody names the field, everybody names the symptom.
Workforce readiness
agent-estimatedSkill-library consolidation is hand-rolled by the lab; no delivery team has run one. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
Servicing agents that learn from resolved cases are the compounding asset; anything learned from customer interactions has to be auditable and removable under APP 12 and CPS 230.
Mechanism · Skill library updated from adjudicated cases only; every entry carries provenance; forgetting on request.
The claims agent in x-continuous-learning is the live case; triage doctrine changes quarterly and a learning agent tracks it without retraining.
Mechanism · Retrieval-updated skills seeded from adjudicated claims; consolidation reviewed by the claims lead.
Merchandising assistants would benefit but the improvement loop is not the blocker there; memory is.
Mechanism · Follows Agentic Memory; no separate mechanism yet.
An agent that learns from citizen interactions raises record-keeping and fairness questions nobody has answered; watch until the audit mechanism exists.
Mechanism · Would need the learned-skill store to be a record under the Archives Act and auditable for drift.
Red team · the strongest case against
The strongest case against: the non-weight route is a memory layer with a different name, and the 11-point gain in the experiment is the agent memorising the claims lead's adjudications, which any RAG over resolved cases would also do. If so the field is a relabelling of Agentic Memory and the 'learning' framing oversells it to clients. The weight-route scepticism may also be stale: a lab shipping production continual learning would collapse the whole field into a model feature.
- —The experiment has no ablation against plain RAG over resolved cases; the gain may be retrieval, not learning.
- —Skill-library contradictions at two months suggest the mechanism does not scale to the timeframes where 'learning' would matter.
- —Two frontier labs are hiring for production continual learning; our 2029 date is an extrapolation from one paper.
Source diversity
- ML research35%
- Practitioners / open-source20%
- Vendor / analyst15%
- Internal / Engel30%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Trigger · The forgetting mechanism named in Agentic Memory's would-have-to-be-true list ships and passes an APP 12 review; without it a learning agent cannot touch client PII and the field cannot move past pilots.
When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.
Trigger · A longitudinal eval exists that measures improvement over deployment weeks on a held-out task family; today every eval is a snapshot, which is why 'does it get better' cannot be answered from the harness.
When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.
If agents learn from use, much of the explicit memory layer becomes the learning substrate rather than a separate product.
A compiled skill wiki is the human-readable form of the skill library; one of the two will absorb the other.
Temporal decision memory is the episodic input a skill library consolidates from.
Weight-level continual learning is one of the capabilities the AGI field is watching for.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- MTMei Tanaka · Research lead · evals2 drops
- TOTom Okafor · Research engineer · agents2 drops
- PRPriya Raman · Research engineer · inference1 drop
- OGOllie Grant · Product engineer1 drop
- ?Anonymous · Anonymous drop1 drop
- RMRohan Mehta · Exec sponsor1 drop
- CDClaire Dubois · Sector owner · Banking1 drop
- MLMarcus Lee · Delivery lead · Telco1 drop
- JPJun Park · Measurement (Nightingale)1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Learning without weights: where continual learning actually lands
strength moderate · 23 citations · review 30 Jun 2027
Non-weight-bound learning via retrieval-updated skills
Retrieval-updated skill pages (the x-personal-wiki mechanism) produce a measurable and persistent behaviour change in an agent without any weight update, and that counts as learning for the purposes of the continuous-learning field.
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Does the 11-point gain survive an ablation against plain RAG over resolved cases?
- 02What contradiction rate in a skill library is the threshold where the gain reverses, and can consolidation hold it under that?
- 03Can a learned skill be attributed to the user who taught it and removed on request, to APP 12 standard?