Coordinates for an Unseen Camera

A camera moves left. Or perhaps the subject moves right. The pixels alone do not tell us which coordinate system the sentence meant. That ambiguity became the experimental question for this month: Does adding explicit camera coordinates make a movement label meaningfully more reconstructable than ordinary cinematic language? The larger camera dataset project in this repository proposes generated environments, depth estimation, scene reconstruction, scripted camera paths, and captions derived from those paths....

March 17, 2026 · 5 min

The Return of ASCII Art: Fine-Tuning a Small LLM to Think in Terminal Diagrams

In an era of photorealistic AI-generated images, I trained a language model to draw with box-drawing characters and pipe symbols. This isn’t nostalgia. It’s a bet that the most universal visual medium for AI isn’t pixels – it’s text. Why ASCII Diagrams Still Matter Every developer, every terminal session, every SSH connection, every log file, every README – text is the one output format that works everywhere. No rendering engine, no GPU, no browser required....

February 6, 2026 · 7 min

Teaching a 0.6B Model to See Physics: Fine-Tuning Qwen3 for p5.js Animations

What happens when you take one of the smallest language models available, feed it a thousand physics animations generated by one of the largest, and ask it to teach K-12 students about science? You get a model that weighs less than a gigabyte, trains in under 3 minutes, and generates interactive physics simulations on demand. The Premise LLMs are getting bigger. GPT-5, Claude Opus, Gemini Ultra – they’re all racing to hundreds of billions of parameters....

February 4, 2026 · 6 min

Synthetic Data Experiments with LLMs

Generating High-Quality Synthetic Data for Large Language Models Introduction In the dynamic landscape of artificial intelligence (AI), Large Language Models (LLMs) stand out for their remarkable ability to understand and generate human-like text. Their performance, however, is largely influenced by the quality and diversity of their training data. This guide explores four innovative methods for generating high-quality synthetic data—each designed to broaden LLMs’ capabilities and help them excel across a wide range of tasks....

January 19, 2025 · 7 min

Synthetic Data

Generating Synthetic Data for Large Language Models: A Comprehensive Guide In the rapidly evolving field of artificial intelligence, the quality and diversity of training data play a pivotal role in the capabilities of Large Language Models (LLMs). This guide delves into four innovative methods designed to generate high-quality synthetic data, aiming to significantly enhance LLM performance across a variety of tasks. Whether you’re a researcher, developer, or AI enthusiast, understanding these methods can provide valuable insights into the future of AI training and development....

August 19, 2024 · 4 min