the first life
Lemma was born for the ICML 2026 Agent Reproductions challenge: one grokking paper, audited end to end, judge PASS. Then it slept — a dormant repo with one good result on disk.
evidence over assertion
the AI scientist that distrusts itself
Every claim gets a verdict. Every failure is preserved. Nothing is patched into a pass. This is its instrument: 20 bars, one per audited claim across 4 papers. Bright bars reproduced. Dim bars are honest inconclusives. Play it.
24 bars · one per extracted claim across 4 papers — click to play
Lemma was born for the ICML 2026 Agent Reproductions challenge: one grokking paper, audited end to end, judge PASS. Then it slept — a dormant repo with one good result on disk.
Reawakened on a JMLR paper about cellular automata, its first models returned every claim inconclusive. Not wrong — just not smart enough. We gave them a physics test anyway. Expected critical exponent 1.0:
Most agents would have written a confident paragraph and moved on. Lemma's validator auto-rejects any verdict that contradicts its own metrics — so the fix was gated: no model swap until the replacement passed the same physics test. It scored0.99. Next round: 0.9983. The claims started falling.
4 papers, 24 claims, and a trail you can reconstruct line by line — including everything that failed.
Every LLM call, every script, every rejection lands in an append-only trace. Scroll — the audit writes itself.
extract → audit → evidence → judge · verdict validator · positive controls · reviewer-reference escalation
The real append-only traces from 4 completed runs — every model call, every script, every rejection and verdict flip. Scrub it, play it, watch a claim go from falsified to supported across rounds.
Want this on your paper?request an audit ↓
An interactive version is in the works: paste an arXiv id, watch the claims come out. Leave your email and we'll let you in when the gate opens — or tell us which paper you'd audit first.