Benchmark rigor and scalability
- structurally held-out geometric evaluator with independently specified latent skills
- random and greedy team baselines
- learning, exploration, initialization, and early/late ablations
- bounded top-k candidate prefilter and request-local scoring caches
- workload-sensitive cost and explicit authorized degraded-route flags
- evidence-gated pair synergy updates
- related-work comparisons, module split, ruff/mypy CI, and release automation
The benchmark remains an author-designed synthetic suite, not evidence of real-world superiority. PyPI upload stays gated until the project owner configures Trusted Publishing; the attached wheel and source distribution passed twine check and clean-environment installation.