Interim citable package while the Zenodo DOI is being finalized.
REFUTE asks whether models can read science honestly. Public writeup: https://bgpt.pro/refute
Current board: 320 questions, 15 complete Truth Scores, leader 74.5 (Grok-4.2). Main finding: critique skill and calibration dissociate.
Assets:
- preprint PDF
- methods notes
- leaderboard JSON
Dataset: https://huggingface.co/datasets/BGPT-OFFICIAL/refute