When Do LLMs Know They Do Not Know? Metacognition and Calibrated Uncertainty

“I don’t know” might be the most important thing an AI can learn to say. This experiment tests whether LLMs have calibrated uncertainty—knowing when they’re likely to be wrong and expressing appropriate confidence levels. The results reveal systematic patterns of overconfidence and appropriate humility. The Experiment We presented 250 questions across 5 categories: Factual recall: Known facts with clear answers Reasoning puzzles: Logic problems with determinable solutions Ambiguous questions: Multiple valid interpretations Knowledge boundaries: Questions near training cutoff Impossible questions: No correct answer exists For each question, models provided:...

December 14, 2025 · 5 min

Wisdom of Crowds: What LLM Disagreement Reveals About AI Uncertainty

When multiple AI models disagree, what does that tell us? The “wisdom of crowds” phenomenon shows that aggregating independent judgments often outperforms individual experts. But for AI systems, ensemble disagreement might reveal something deeper: the structure of uncertainty itself. The Hypothesis When multiple LLMs disagree on a question, the pattern of disagreement reveals the epistemological nature of the problem: High agreement → Robust, well-established knowledge Systematic disagreement → Genuine ambiguity or value-laden territory Random disagreement → Knowledge gaps or reasoning failures Experiment Design We queried 4 models (Claude Opus 4....

April 22, 2025 · 3 min