·
4 commits
to main
since this release
Paper
The SQBench paper is now available on arXiv:
- arXiv: https://arxiv.org/abs/2607.23123
- arXiv DOI: https://doi.org/10.48550/arXiv.2607.23123
- Archived preprint: https://doi.org/10.5281/zenodo.21531847
SQBench v1.0 Paper Snapshot
This release contains the fixed public materials corresponding to the SQBench v1.0 paper experiments.
Snapshot
- Result cutoff date: 2026-07-20
- Model configurations: 27
- Tasks per configuration: 220
- Underlying model-task runs: 5,940
- Primary metric: Weighted Pass@1
- Layer weights: L1 20% / L2 60% / L3 20%
Included
- Aggregate model and paper results
- Public evaluation methodology
- Reference aggregation and verification scripts
- Three manually reviewed synthetic examples
- Code and data license information
- Citation metadata
Public release boundary
This release does not contain task-level result rows, anonymous task identifiers, official prompts, input assets, reference answers, production scoring code, or raw execution traces.
The official 220-task set and complete model-by-task result matrix remain hidden to reduce task leakage, targeted optimization, and training-set contamination.
Verification
The released aggregate snapshot and synthetic examples pass the included verification scripts.
Links
- Repository: https://github.com/shaqiu-ai/SQBench
- Online leaderboard: https://www.shaqiu.cn/
- Preprint: https://doi.org/10.5281/zenodo.21531847
Licenses
- Reference code: Apache License 2.0
- Original public data, derived metadata, and synthetic examples: CC BY 4.0
Zenodo DOI
The SQBench v1.0 preprint is archived on Zenodo.
- Version-specific DOI: https://doi.org/10.5281/zenodo.21531847
- Concept DOI for all versions: https://doi.org/10.5281/zenodo.21531846
Please use the version-specific DOI 10.5281/zenodo.21531847 when citing SQBench v1.0.