What I Learned Running a Long Horizon Memory Experiment on 4 A100 GPUs
I wanted to answer one practical question. Can a model keep learning over long sessions without slowly losing grip on earlier facts? This post is a learning oriented walkthrough of one real campaign I ran. It focuses on understanding and decision making, not just reporting scores. Code and implementation are here: GitHub repo: rlm-experiment-codex Live report dashboard Overview I compared two memory methods with the same base model and the same datasets....