Deep dives

Technical writing on independent verification, the claim-vs-effect gap, and the falsification record.

the essays, written to be checked
case-studyagentshermesself-evolution September 18, 2026

Hermes Claims Skills Cut Tokens 30-85%. They Added Tokens on All 3 Tests.

We ran the most-used AI agent of 2026 through 8 tasks, a memory test, and 3 token comparisons. The tasks and memory held. The token claim didn't.

case-studyfalsificationcode September 17, 2026

An AI Said It Fixed SymPy #24909. It Didn't.

A real bug, a non-empty patch, a confident 'fixed,' and a test that still fails. The full falsification, with the signed receipt you can check.

verificationclaim-vs-effectfalsification September 17, 2026

The Claim-vs-Effect Gap

Why an AI system's report of its own success is a claim, not evidence, and how to tell which one you're looking at.