Skip to content

DiamondBench v1.1: auditable AI gemology benchmark

Latest

Choose a tag to compare

@JacobiusMakes JacobiusMakes released this 04 Sep 20:07
· 2 commits to main since this release

DiamondBench v1.1 packages the 47-question benchmark, sourced answer key,
deterministic grader, raw model answers, and dated results into a fixed release.

Two search-grounded Gemini 2.5 Flash runs, seven weeks apart, each scored 45 of
47 questions as pass and 2 as partial, or 95.7%. The same two provenance gaps
appeared in both runs. No answer in either run was graded as false.

The initial September scoring pass exposed five grader errors. The rules were
corrected, every September answer was regraded, and the July run was regraded
under the same current rules. The July result did not change. The repository
includes the raw answers, the regrade script, methodology, and known limits so
the result can be audited.

Release assets include the machine-readable question set, Hugging Face JSONL,
September raw results, human-readable scoreboard, and chart. The tagged source
also includes citation metadata and the complete runnable benchmark.