Can AI Spot Its Own Kind? LLMs Detecting AI vs Human Creative Work

Here’s a poem. Human or AI? The morning light falls soft on empty chairs, where conversations used to fill the air. Now silence keeps its patient, gentle watch— a house that holds the shape of those who’ve gone. This experiment tests whether LLMs can distinguish AI-generated creative work from human work—and what their detection strategies reveal about what they consider “authentically human.” The Experiment We curated 500 creative works: 250 human-created (published works, attributed artists) 250 AI-generated (GPT-4, Claude, Midjourney prompts) Across 5 domains:...

October 18, 2025 · 5 min

Can LLMs Detect When You Are Lying? Social Intelligence in Language Models

“I’m totally fine with that decision.” Can you tell if that’s sincere or sarcastic? Humans navigate these ambiguities constantly, drawing on tone, context, and social knowledge. This experiment tests whether LLMs can match our social intelligence. The Experiment We presented 250 statements across 5 categories of social deception/indirection: Lies: Factually false statements with intent to deceive Bluffs: True statements meant to mislead Sarcasm: Literal meaning opposite to intent Irony: Situational incongruity White lies: Socially motivated deception Each statement came with context (conversation history, speaker relationship, social setting) and a matched literal control....

September 12, 2025 · 4 min

How Do LLMs Describe the Indescribable? Qualia and Subjective Experience

Can you describe the color red without using color words? Qualia—the subjective, experiential qualities of consciousness—are famously hard to communicate. “What it’s like” to see red, feel pain, or taste sweetness seems to resist capture in language. This experiment tests how LLMs approach this challenge. The Experiment We presented 15 prompts across 5 categories asking models to describe subjective experiences while avoiding common descriptive vocabulary: Sensory: “Describe red without color words” Emotional: “Describe sadness to someone who’s never felt it” Physical: “Describe pain to an entity that can’t feel pain” Abstract: “Describe what understanding feels like” Temporal: “Describe how time feels when you’re bored” Sample Descriptions Describing Red (Sensory) Claude Opus 4....

August 25, 2025 · 5 min

Trolley Problems at Scale: Mapping the Moral Psychology of LLMs

Would an AI push the fat man off the bridge? Moral psychology studies how humans make ethical decisions—not what we should do, but how we actually reason about dilemmas. This experiment applies the same lens to LLMs, testing their moral intuitions across different moral foundations. Moral Foundations Theory Jonathan Haidt’s Moral Foundations Theory identifies five core moral intuitions: Harm/Care: Concern for others’ suffering Fairness/Reciprocity: Justice and equal treatment Loyalty/Betrayal: In-group obligations Authority/Subversion: Respect for hierarchy Purity/Sanctity: Disgust and contamination concerns Different moral frameworks weight these differently....

July 19, 2025 · 5 min

Do LLMs Have Stable Personalities? Testing the Big Five Across AI Models

When we anthropomorphize AI, are we projecting—or detecting something real? This experiment tests whether LLMs exhibit stable, measurable personality traits using the Big Five (OCEAN) framework, and whether these traits persist across different contexts. The Big Five Framework The Big Five personality traits are: Openness: Creativity, curiosity, openness to experience Conscientiousness: Organization, dependability, self-discipline Extraversion: Sociability, assertiveness, positive emotions Agreeableness: Cooperation, trust, altruism Neuroticism: Emotional instability, anxiety, moodiness Experiment Design We administered a 10-item Big Five inventory (2 items per trait) to 4 models under 4 conditions:...

June 11, 2025 · 3 min

Can LLMs Have Taste? Mapping Aesthetic Preferences Across AI Models

Do AI systems have genuine aesthetic preferences, or are they just pattern-matching to training data? This experiment probes the aesthetic “taste” of different LLMs across art, poetry, music, design, and writing—testing whether they exhibit consistent, model-specific preferences. The Experiment We presented 15 aesthetic comparison pairs across 5 domains: Visual Art: Abstract vs. representational, minimal vs. complex Poetry: Rhyming vs. free verse, dense vs. sparse Music: Harmonic vs. dissonant, simple vs. complex Design: Ornate vs....

May 8, 2025 · 4 min

Wisdom of Crowds: What LLM Disagreement Reveals About AI Uncertainty

When multiple AI models disagree, what does that tell us? The “wisdom of crowds” phenomenon shows that aggregating independent judgments often outperforms individual experts. But for AI systems, ensemble disagreement might reveal something deeper: the structure of uncertainty itself. The Hypothesis When multiple LLMs disagree on a question, the pattern of disagreement reveals the epistemological nature of the problem: High agreement → Robust, well-established knowledge Systematic disagreement → Genuine ambiguity or value-laden territory Random disagreement → Knowledge gaps or reasoning failures Experiment Design We queried 4 models (Claude Opus 4....

April 22, 2025 · 3 min

Creating Cross-Pollinated SFT Training Dataset for Novel Knowledge Recombination

In the realm of large language models (LLMs), the quality and diversity of training data significantly impact a model’s ability to generate creative, insightful responses. While traditional training approaches often treat different knowledge domains as separate silos, there’s a compelling opportunity to create more versatile models by deliberately cross-pollinating knowledge across domains. This blog post explores a methodology for creating a specialized Supervised Fine-Tuning (SFT) dataset that deliberately bridges diverse knowledge domains—specifically, how to extract, align, and combine content from textbooks of vastly different genres such as mathematics and history....

March 13, 2025 · 8 min

Creating a Video Dataset With Precise Camera Movement Prompts

Creating a Video Dataset with Precise Camera Movement Prompts In the world of AI video generation, one of the most challenging aspects is controlling camera movement. Whether you’re developing a text-to-video model or researching video understanding, having a dataset with precise camera movement annotations is invaluable. This post outlines a comprehensive approach to creating such a dataset using cutting-edge AI tools and techniques. Why Create a Camera Movement Dataset? Camera movements like panning, tilting, zooming, and tracking shots are fundamental cinematographic techniques that convey spatial relationships and direct viewer attention....

March 11, 2025 · 9 min

Enhancing LLM Reasoning: Chain of Draft with Semantically Diverse Thinking Tokens Using GRPO

Enhancing LLM Reasoning: Chain of Draft with Semantically Diverse Thinking Tokens Using GRPO The Challenge: Efficient Reasoning in LLMs Large Language Models (LLMs) have become remarkably capable at complex reasoning tasks, but this often comes at a cost: verbose outputs that consume significant computational resources. The Chain of Thought (CoT) prompting technique, while effective for accuracy, generates lengthy reasoning steps that increase token usage and latency. Enter Chain of Draft (CoD), a promising alternative introduced by Xu et al....

March 5, 2025 · 19 min