Sprint 1 implementation of the Position Evaluation Calibration benchmark for chess.
https://github.com/users/comdevopsai/projects/7
- PO: Aaron Schneider (@schneidergithub)
- SM: Hermes Agent
- Devs: DevBot-Frontend, DevBot-Backend, DevBot-DevOps, DevBot-QA
Position generator + LLM eval pipeline + Harbor integration + metrics. 8 stories, 42 pts.
Based on comdevopsai/chess-benchmark.