We tried to break our own AI's brain — it rediscovered CRISPR, AlphaFold 2, and mRNA vaccines from scratch (8/9, receipts included) #40
UditAkhourii
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hello @everyone 👋
Evals for ADHD Phase #2 are here and tbh they are nothing less than shocking.
v0.1 got the obvious objection: same-model judge, 6 problems, engineering only. So Phase 2 tried to break our own result — new judge model, new domains, and the sharpest test we could design: a ground-truth eval where the answer is a real published discovery, not another LLM's opinion.
The big one: we took 9 real breakthroughs — CRISPR, AlphaFold 2, mRNA vaccine delivery, Raft, GLP-1's effect on addiction pathways — stripped every name and term that could let a model pattern-match, and gave ADHD only what was known before each discovery. Did the candidate pool contain the real mechanism?
8 of 9. 5 of 5 on everything post-training-cutoff, where memorization isn't on the table. 🤯
Best detail in the whole phase: on several hits, the right answer was sitting in the pool at rank 12–25 of 30 — the divergent branches found it, and the critic pass nearly buried it. Real, fixable weak point, going straight into v0.2.
Also ran 12 engineering problems under a second judge model, plus three new domains (strategy, health, biochemistry) to test if this holds outside code. All four studies live now.
Everything's public, including raw logs — not just summary tables:
📄 Full paper + reports:
findings/💾 Raw per-call logs:
research/logs/📊 Results JSON, if you want to re-crunch it yourself
🔗 https://github.com/DivergentLab/evals-for-adhd
Straight talk: one health case (metformin) has a self-contradicting writeup — reasoning says MISS, label says HIT. Fixing this week. Study 4's frame-ablation also has a broken metric (every frame showing 0% survival, which doesn't square with Study 3) — flagging it, not hiding it. Dig in if you want.
Small N (9 cases, not 900) — strong early signal, not a settled law. But it's checkable against the literature instead of an LLM's opinion, which is why it's the one we most want you all to pressure-test.
Go break it. 🔨
All reactions