I wanted to answer one practical question.
Can a model keep learning over long sessions without slowly losing grip on earlier facts?
This post is a learning oriented walkthrough of one real campaign I ran. It focuses on understanding and decision making, not just reporting scores.
Code and implementation are here:
Overview
I compared two memory methods with the same base model and the same datasets.
- Text Buffer
- Latent Pager Memory, or LPM
Setup:
- Model: Qwen3 1.7B
- Hardware: 4x A100 80GB
- Data: oolong real and oolong synth
- Focus: continuous learning behavior, context rot resistance, long horizon stability
I used three benchmark families.
- Sentinel checks for baseline quality and contradiction
- Context rot stress for distractor pressure and horizon growth
- Continual stream tests for memory performance across long event sequences
Architecture in plain language
The pipeline is simple to explain.
- Build prepared examples from both datasets.
- Run the same model with each memory method.
- Score each run with the same metric stack.
- Save per example records and summary files.
- Render a dashboard and a static report.
The most important operational lesson came late. Top level parallelization was good at first, but heavy tail jobs created idle GPUs at the end. I fixed that by sharding the remaining hard context rot conditions across all four GPUs.
Results
Sentinel baseline


| Dataset | Method | Task score | Contradiction | Hall total | Mean time s | Mean total tokens |
|---|---|---|---|---|---|---|
| oolong real | LPM | 0.0669 | 0.0819 | 0.0028 | 1.945 | 654.6 |
| oolong real | Text Buffer | 0.0515 | 0.1040 | 0.0111 | 46.511 | 55337.1 |
| oolong synth | LPM | 0.2878 | 0.1753 | 0.3694 | 1.664 | 459.3 |
| oolong synth | Text Buffer | 0.2500 | 0.2055 | 0.3639 | 4.335 | 2469.8 |
What this taught me:
- LPM improved task score in both datasets.
- LPM reduced contradiction in both datasets.
- Latency difference on oolong real was very large.
Context rot behavior


How to read this section:
- Distractor ratio increases retrieval pressure.
- Horizon multiplier increases memory distance.
- Falling task score in harder cells means context rot sensitivity.
- Missing cells mean active jobs were still running at snapshot time.
What I learned from context rot so far:
- Some LPM regions are stable even at higher stress.
- Text Buffer has a heavier runtime tail in hard real dataset conditions.
- Hard corner data should be interpreted only after all shards finish.
Continual stream behavior

In this event budget, capacity effects were smaller than I expected. That suggests the current stream was not yet harsh enough to separate methods strongly by memory size alone.
Training setup and why it matters
This campaign did not update model weights.
- Base weights stayed frozen.
- Memory behavior was the thing under test.
- Seed and split were fixed for comparability.
This matters for learning because it isolates memory system quality from optimizer noise.
Ablations
The cleanest ablation was method swap with everything else fixed.
| Dataset | LPM minus Text Buffer task delta | LPM minus Text Buffer contradiction delta | LPM speedup |
|---|---|---|---|
| oolong real | +0.0154 | -0.0221 | 23.92x |
| oolong synth | +0.0378 | -0.0302 | 2.60x |
This is the exact ablation I care about most in production planning because it combines quality and cost in the same direction.
Visual timeline of the run

Timeline insight:
- Sentinel finished early.
- Continual stream finished next.
- Context rot hard shards dominated tail time.
- Tail sharding recovered utilization.
Hypotheses and what changed in my thinking
I started with three hypotheses.
- LPM should improve cost while preserving or improving quality.
- Context rot should worsen as distractors and horizon increase.
- Continual memory should keep high hit rate over long sessions.
Current interpretation:
- Hypothesis one looks strong.
- Hypothesis two is supported in trend, with final hard corner pending.
- Hypothesis three looks promising but needs longer stream stress to verify limit behavior.
Examples from outputs
Real per example records matter because averages can hide failure modes.
- Several count questions were answered exactly.
- Some hard count queries still returned unknown.
- Context classification could succeed on one item and fail on a nearby item in the same condition.
This pattern is a reminder that reliability is about tails, not only means.
Future direction for RLM models
I want this section to be honest. The chart below is a forecast, not measured data.

Forecast assumptions:
- Better memory routing and retrieval gating continue to improve retention.
- Contradiction controls become first class objectives in memory systems.
- Cost efficiency improves as latent memory representations mature.
Prediction table:
| Year | What likely improves | What will still be hard |
|---|---|---|
| 2026 | Better memory routing, lower latency variance | Very long context factual consistency |
| 2027 | More stable contradiction control in long dialogs | Generalization under domain shift |
| 2028 | Memory systems become default in production agent stacks | Evaluation of rare tail failures at scale |
My current belief is that the winning RLM direction is not one trick. It is a blend of better memory representation, better retrieval policies, and better stress evaluation loops that run continuously.
What to do next if you are building this yourself
- Keep a sentinel suite running all the time.
- Add explicit context rot stress tests early.
- Track tail latency and tail correctness, not only mean values.
- Keep your reporting pipeline automatic so you can learn every day, not at the end.
This project taught me that long horizon capability is not a single benchmark property. It is an engineering discipline across memory design, evaluation design, and runtime scheduling.