<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Synthetic Data on Dylan Ler</title><link>http://dylanler.github.io/tags/synthetic-data/</link><description>Recent content in Synthetic Data on Dylan Ler</description><generator>Hugo -- 0.133.0</generator><language>en-us</language><lastBuildDate>Thu, 27 Aug 2026 10:00:00 -0700</lastBuildDate><atom:link href="http://dylanler.github.io/tags/synthetic-data/index.xml" rel="self" type="application/rss+xml"/><item><title>Coordinates for an Unseen Camera</title><link>http://dylanler.github.io/posts/coordinates-for-an-unseen-camera/</link><pubDate>Tue, 17 Mar 2026 07:42:00 -0700</pubDate><guid>http://dylanler.github.io/posts/coordinates-for-an-unseen-camera/</guid><description>A camera moves left. Or perhaps the subject moves right. The pixels alone do not tell us which coordinate system the sentence meant.
That ambiguity became the experimental question for this month:
Does adding explicit camera coordinates make a movement label meaningfully more reconstructable than ordinary cinematic language?
The larger camera dataset project in this repository proposes generated environments, depth estimation, scene reconstruction, scripted camera paths, and captions derived from those paths.</description></item><item><title>The Return of ASCII Art: Fine-Tuning a Small LLM to Think in Terminal Diagrams</title><link>http://dylanler.github.io/posts/ascii-tui-diagrams-fine-tuning-qwen3/</link><pubDate>Fri, 06 Feb 2026 08:00:00 -0800</pubDate><guid>http://dylanler.github.io/posts/ascii-tui-diagrams-fine-tuning-qwen3/</guid><description>In an era of photorealistic AI-generated images, I trained a language model to draw with box-drawing characters and pipe symbols.
This isn&amp;rsquo;t nostalgia. It&amp;rsquo;s a bet that the most universal visual medium for AI isn&amp;rsquo;t pixels &amp;ndash; it&amp;rsquo;s text.
Why ASCII Diagrams Still Matter Every developer, every terminal session, every SSH connection, every log file, every README &amp;ndash; text is the one output format that works everywhere. No rendering engine, no GPU, no browser required.</description></item><item><title>Teaching a 0.6B Model to See Physics: Fine-Tuning Qwen3 for p5.js Animations</title><link>http://dylanler.github.io/posts/fine-tuning-qwen3-p5js-physics-animations/</link><pubDate>Wed, 04 Feb 2026 09:30:00 -0800</pubDate><guid>http://dylanler.github.io/posts/fine-tuning-qwen3-p5js-physics-animations/</guid><description>What happens when you take one of the smallest language models available, feed it a thousand physics animations generated by one of the largest, and ask it to teach K-12 students about science?
You get a model that weighs less than a gigabyte, trains in under 3 minutes, and generates interactive physics simulations on demand.
The Premise LLMs are getting bigger. GPT-5, Claude Opus, Gemini Ultra &amp;ndash; they&amp;rsquo;re all racing to hundreds of billions of parameters.</description></item><item><title>Synthetic Data Experiments with LLMs</title><link>http://dylanler.github.io/posts/synthetic-data-experiments/</link><pubDate>Sun, 19 Jan 2025 20:45:48 -0800</pubDate><guid>http://dylanler.github.io/posts/synthetic-data-experiments/</guid><description>A comprehensive guide exploring four innovative methods for generating high-quality synthetic data for Large Language Models, including persona-driven web crawling, graph-based reasoning, research paper extraction, and curriculum learning.</description></item><item><title>Synthetic Data</title><link>http://dylanler.github.io/posts/synthetic-data/</link><pubDate>Mon, 19 Aug 2024 02:19:26 -0700</pubDate><guid>http://dylanler.github.io/posts/synthetic-data/</guid><description>Generating Synthetic Data for Large Language Models: A Comprehensive Guide In the rapidly evolving field of artificial intelligence, the quality and diversity of training data play a pivotal role in the capabilities of Large Language Models (LLMs). This guide delves into four innovative methods designed to generate high-quality synthetic data, aiming to significantly enhance LLM performance across a variety of tasks. Whether you&amp;rsquo;re a researcher, developer, or AI enthusiast, understanding these methods can provide valuable insights into the future of AI training and development.</description></item></channel></rss>