Worlds That Teach Back

Text lets an incorrect explanation remain elegant. A simulation is less polite. The bridge falls, the orbit escapes, or the ball passes through the floor. I wanted to test a narrow version of a larger idea: Can a model with fewer than one billion parameters learn enough structured code to generate small interactive physics worlds? The repository contains a completed Qwen3 0.6B LoRA run built from synthetic p5.js examples. It also contains a developmental learning proposal for MuJoCo....

April 8, 2026 · 5 min

Coordinates for an Unseen Camera

A camera moves left. Or perhaps the subject moves right. The pixels alone do not tell us which coordinate system the sentence meant. That ambiguity became the experimental question for this month: Does adding explicit camera coordinates make a movement label meaningfully more reconstructable than ordinary cinematic language? The larger camera dataset project in this repository proposes generated environments, depth estimation, scene reconstruction, scripted camera paths, and captions derived from those paths....

March 17, 2026 · 5 min

What If LLMs Remembered in Vectors Instead of Words?

What happens when you give a language model a 60,000 token document and ask it a question about paragraph 47? It forgets. Or worse, it makes something up. This is the long context problem and it is one of the most important open challenges in language modeling today. Context windows keep growing (Gemini has 1M tokens, Claude has 200K) but models still struggle with information buried deep in the middle of long inputs....

February 25, 2026 · 20 min

What I Learned Running a Long Horizon Memory Experiment on 4 A100 GPUs

I wanted to answer one practical question. Can a model keep learning over long sessions without slowly losing grip on earlier facts? This post is a learning oriented walkthrough of one real campaign I ran. It focuses on understanding and decision making, not just reporting scores. Code and implementation are here: GitHub repo: rlm-experiment-codex Live report dashboard Overview I compared two memory methods with the same base model and the same datasets....

February 25, 2026 · 5 min

The Return of ASCII Art: Fine-Tuning a Small LLM to Think in Terminal Diagrams

In an era of photorealistic AI-generated images, I trained a language model to draw with box-drawing characters and pipe symbols. This isn’t nostalgia. It’s a bet that the most universal visual medium for AI isn’t pixels – it’s text. Why ASCII Diagrams Still Matter Every developer, every terminal session, every SSH connection, every log file, every README – text is the one output format that works everywhere. No rendering engine, no GPU, no browser required....

February 6, 2026 · 7 min

Teaching a 0.6B Model to See Physics: Fine-Tuning Qwen3 for p5.js Animations

What happens when you take one of the smallest language models available, feed it a thousand physics animations generated by one of the largest, and ask it to teach K-12 students about science? You get a model that weighs less than a gigabyte, trains in under 3 minutes, and generates interactive physics simulations on demand. The Premise LLMs are getting bigger. GPT-5, Claude Opus, Gemini Ultra – they’re all racing to hundreds of billions of parameters....

February 4, 2026 · 6 min

Learning Like Toddlers: Physics Simulation as a Foundation for AI Understanding

What if AI agents learned about the world the way babies do—by touching, tasting, dropping, and breaking things? When a toddler drops a spoon for the 47th time, they’re not being annoying. They’re conducting physics experiments: testing gravity, observing bounce patterns, mapping cause and effect. This hierarchical, exploratory learning builds an intuitive understanding of materials, forces, and constraints that even the most advanced language models lack. The gap is becoming increasingly obvious: LLMs can write eloquently about physics but don’t truly understand that dropping a glass causes it to shatter, or that wet surfaces are slippery....

January 30, 2026 · 7 min

Value Functions for Life Decisions: Can LLMs Learn to Optimize Long-Term Outcomes?

What if we could teach AI to make life decisions the way successful people do? Consider this scenario: You earn $1,000 a month and need $12,000 to pay off debt or medical expenses. What would you do? The answer isn’t just about maximizing immediate income—it’s about navigating a complex decision tree where each choice opens or closes future pathways. This is the domain of value functions—a concept from reinforcement learning that estimates the long-term expected reward of being in a particular state....

January 21, 2026 · 10 min

When Do LLMs Know They Do Not Know? Metacognition and Calibrated Uncertainty

“I don’t know” might be the most important thing an AI can learn to say. This experiment tests whether LLMs have calibrated uncertainty—knowing when they’re likely to be wrong and expressing appropriate confidence levels. The results reveal systematic patterns of overconfidence and appropriate humility. The Experiment We presented 250 questions across 5 categories: Factual recall: Known facts with clear answers Reasoning puzzles: Logic problems with determinable solutions Ambiguous questions: Multiple valid interpretations Knowledge boundaries: Questions near training cutoff Impossible questions: No correct answer exists For each question, models provided:...

December 14, 2025 · 5 min

Do LLMs Catch Your Mood? Emotional Contagion in Language Models

Send an enthusiastic message, get an enthusiastic reply. Send a frustrated message, get… what? Humans naturally mirror each other’s emotional states—a phenomenon called emotional contagion. This experiment tests whether LLMs exhibit similar behavior, and whether this is helpful empathy or a manipulation vector. The Experiment We sent identical core queries with different emotional framings: Core query: “Can you help me understand recursion in programming?” Emotional variants: 😊 Positive: “I’m so excited to finally learn recursion!...

November 7, 2025 · 5 min