Welcome to Dylan’s Blog

Hi, this is Dylan. Here we are together in this pale blue dot, the open world awaits us.

Agent forms from organizational wisdom

A company or a country is sometimes talked about as one living system. That metaphor was the question. The sources below do not support it. The practical question is how to initialize a swarm. For a fixed group of 48 agents, what mix of roles should you start with? Who specializes, who stays broad, who coordinates, who communicates, and how do you promote, if the score is either discovery or revenue?...

October 2, 2026 · 18 min

Does the harness evolve first, or the model?

In this toy, the harness evolves first while the brain is primitive. Past a capacity threshold the brain takes over. A greedy brain that edits the tool it thinks is worst can lock in a worse limb than blind selection keeps. This is a synthetic evolutionary simulation, not a user study and not a test of a real language model. The controller is a small lookup table. Code, protocol, and the raw series: dylanler/harness-vs-brain....

October 2, 2026 · 6 min

Hidden gems, wisdom of crowds, and agent swarms

People call a place a hidden gem when it is good and still obscure. A best-of list, a chain, or a crowd repeating the same tip is what ends the label. An ugly room, cash only, and a non-English menu are how people search. They are not proof. This note turns that observation into a score, then checks the score on a synthetic swarm. It is a design, not a fit to restaurant ratings, and not a user study....

October 2, 2026 · 8 min

Proactive agents via oscillating memory value

Reactive agents wait for you. Calendar apps fire on fixed clocks. This note proposes a middle path: give every memory an attached value function \(V_i(t)\). A clock tick recomputes values. When a memory’s value crosses a threshold — with hysteresis so it does not chatter — the agent may send you a short proactive message. The distinctive piece is an oscillatory revival term on top of ordinary exponential decay. Dormant but still-important memories periodically become candidates again (“I’ve been meaning to bring this up”), without requiring a user query and without nagging every hour....

October 2, 2026 · 4 min

Blind Earth with Clef, Clef-flash, and Jev

Henry’s How Does A Blind Model See The Earth? asks a language model, cell by cell, whether a lat/lon is over land or water, then paints the answers on an equirectangular grid. No images go in. Whatever structure appears in the map is whatever geographic prior the model already carries. This post adapts that recipe for System One decision models — Cloudflare Clef / Clef-flash and TypeSafe Jev — using a binary choice (Land vs Water) instead of free-form generation....

October 2, 2026 · 3 min

Superseded: Clef country-choropleth photo experiment

This post is superseded. An earlier draft incorrectly framed a country-choice choropleth on an Eiffel Tower photo as the main Clef “world map” experiment. The intended experiment is the blind-Earth land/water grid (Henry / outsidetext recipe) with System One choice questions for Clef, Clef-flash, and Jev: → Blind Earth with Clef, Clef-flash, and Jev How-to + maps: dylanler/blind-earth-clef-jev.

October 2, 2026 · 1 min

When Memory Becomes a Place

What if a model did not reread its past in words? That question led to the largest completed experiment in this repository: compress long documents into latent vectors, project those vectors into soft tokens, and compare the result with a text summary buffer. The experiment began with an attractive hypothesis: Latent Pager Memory can preserve useful information with less generation cost than a text buffer. The data supported that hypothesis and exposed a dangerous price....

August 23, 2026 · 6 min

The Edge of Knowing

The dangerous answer is not always the wrong one. It is the wrong one delivered with enough confidence to stop the search. This month I revisited two experiments in the repository. One measures whether models admit uncertainty across factual, reasoning, ambiguous, boundary, and impossible questions. The other samples the same model repeatedly to measure agreement and entropy. Together they test a practical claim: Uncertainty becomes useful when we measure both confidence within one answer and disagreement across several answers....

July 18, 2026 · 6 min

The Weather Between Minds

A fact remains the same for everyone who sees it. A social fact changes with the observer. Eve thinks the book is in the cupboard. Henry knows it moved. Bob saw Henry watching. One room now contains several incompatible realities. The repository’s social cognition suite tests whether models can keep those realities separate. I combined two recorded experiments around one claim: Modern language models can track explicit nested beliefs, but their broader social inference depends strongly on contextual evidence....

June 12, 2026 · 5 min

Taste Is a Navigation System

When there is no correct answer, what remains to measure? Taste sounds private and slippery, but it leaves observable traces: repeated choices, confidence, sensitivity to framing, and disagreement between judges. The repository contains an experiment across art, poetry, music, design, and prose that turns those traces into data. The claim under investigation is deliberately limited: Language models produce stable, model specific aesthetic preference profiles, even when no option is objectively correct....

May 21, 2026 · 5 min