Proposal
Add REFUTE to related / evaluation / research tooling docs if this project surfaces LLM evaluation, RAG quality, or scientific agent workflows.
REFUTE is a scientific critique + calibration benchmark (paper-grounded claims → predictions → judge scores → Brier/ECE).
Happy to adjust wording / PR if preferred.
Proposal
Add REFUTE to related / evaluation / research tooling docs if this project surfaces LLM evaluation, RAG quality, or scientific agent workflows.
REFUTE is a scientific critique + calibration benchmark (paper-grounded claims → predictions → judge scores → Brier/ECE).
Happy to adjust wording / PR if preferred.