Repository navigation
Third release of deepwolf — five substantial features since v0.2.0.
Added
- Copilot calibration —
deepwolf calibratescores the copilot's own probabilities (Brier score, skill score, Murphy decomposition, reliability diagram), so a human knows how much to trust its percentages. Motivated by the WOLF benchmark. - Bidding-based discussion (#19) — agents bid for the discussion floor (
--bidding); bids are public, so wanting to talk is itself a tell. Inspired by Google's Werewolf Arena. - Arena leaderboard (#4) —
deepwolf leaderboardranks agents fairly under identical seeds. - Simplified Chinese support — the whole game runs in English or Chinese (
--lang zh); addsREADME.zh-CN.md. - Provider guide (#5) —
docs/providers.mdfor connecting any OpenAI-compatible endpoint.
81 tests, green CI on Python 3.10-3.12. English output is byte-for-byte unchanged from v0.2.0.
pip install -e .
deepwolf simulate --players 8 --seed 1 --lang zh --bidding
deepwolf calibrate --games 60