Skip to content

Releases: JuneQQQ/deepwolf

deepwolf 0.3.0

Choose a tag to compare

@JuneQQQ JuneQQQ released this 18 May 17:12

Third release of deepwolf — five substantial features since v0.2.0.

Added

  • Copilot calibration — deepwolf calibrate scores the copilot's own probabilities (Brier score, skill score, Murphy decomposition, reliability diagram), so a human knows how much to trust its percentages. Motivated by the WOLF benchmark.
  • Bidding-based discussion (#19) — agents bid for the discussion floor (--bidding); bids are public, so wanting to talk is itself a tell. Inspired by Google's Werewolf Arena.
  • Arena leaderboard (#4) — deepwolf leaderboard ranks agents fairly under identical seeds.
  • Simplified Chinese support — the whole game runs in English or Chinese (--lang zh); adds README.zh-CN.md.
  • Provider guide (#5) — docs/providers.md for connecting any OpenAI-compatible endpoint.

81 tests, green CI on Python 3.10-3.12. English output is byte-for-byte unchanged from v0.2.0.

pip install -e .
deepwolf simulate --players 8 --seed 1 --lang zh --bidding
deepwolf calibrate --games 60

deepwolf 0.2.0

Choose a tag to compare

@JuneQQQ JuneQQQ released this 18 May 11:10

Second release of deepwolf — three roadmap features land.

Added

  • Hunter role (#1) — on death (lynched or killed at night) the Hunter takes one living player down with them; chained Hunter deaths resolve correctly.
  • Witch role (#2) — two one-time potions: each night the Witch learns the werewolves' victim and may heal them and/or poison any player. The night phase now resolves multiple simultaneous deaths.
  • JSON transcript export (#3) — deepwolf.game.transcript and a --transcript PATH flag write a finished game as a versioned, machine-readable record.

43 tests, green CI on Python 3.10–3.12. Five roles now in play: Villager, Werewolf, Seer, Doctor, Hunter, Witch.

pip install -e .
deepwolf simulate --players 8 --seed 1

deepwolf 0.1.0

Choose a tag to compare

@JuneQQQ JuneQQQ released this 18 May 10:50

First public release of deepwolf — an LLM werewolf engine.

Highlights

  • Self-play arena — LLM agents play full games of werewolf so you can benchmark reasoning, deception and deduction under hidden information.
  • Human copilot — an explainable werewolf-suspicion model that advises your vote while you play.
  • Strict seeded rules engine — Villager / Werewolf / Seer / Doctor, fully reproducible games, illegal agent moves validated away.
  • Vendor-neutral LLMs — any OpenAI-compatible endpoint; ships with a deterministic offline mock so everything runs with zero setup.
  • CLI — deepwolf simulate, deepwolf arena, deepwolf play.

33 tests, CI across Python 3.10–3.12. See the README to get started in one command:

deepwolf simulate --players 7 --seed 1