Skip to content

deepwolf 0.3.0

Latest

Choose a tag to compare

@JuneQQQ JuneQQQ released this 18 May 17:12
· 3 commits to main since this release

Third release of deepwolf — five substantial features since v0.2.0.

Added

  • Copilot calibration — deepwolf calibrate scores the copilot's own probabilities (Brier score, skill score, Murphy decomposition, reliability diagram), so a human knows how much to trust its percentages. Motivated by the WOLF benchmark.
  • Bidding-based discussion (#19) — agents bid for the discussion floor (--bidding); bids are public, so wanting to talk is itself a tell. Inspired by Google's Werewolf Arena.
  • Arena leaderboard (#4) — deepwolf leaderboard ranks agents fairly under identical seeds.
  • Simplified Chinese support — the whole game runs in English or Chinese (--lang zh); adds README.zh-CN.md.
  • Provider guide (#5) — docs/providers.md for connecting any OpenAI-compatible endpoint.

81 tests, green CI on Python 3.10-3.12. English output is byte-for-byte unchanged from v0.2.0.

pip install -e .
deepwolf simulate --players 8 --seed 1 --lang zh --bidding
deepwolf calibrate --games 60