v0.2.0
Feature release. Expands the v0.1.0 evaluation (4 models, 8 conditions) into a layered suite with independent control of action protocol, communication structure, group size, and prompting strategy.
Added
- Three-layer experiment suite (
experiments/configs/conditions.yaml):
Layer 1 universality (six baseline conditions: simultaneous/sequential × communication/no-communication × N ∈ {5, 10}, run across all models),
Layer 2 scaling (extends group size to N ∈ {7, 15, 20}),
Layer 3 ablations. - Ablations: prompt variant (symmetry breaking, theory of mind, resource ordering, minimal), memory window (k ∈ {3, 5}), multi-round communication (3, 5 rounds), temperature (0.3, 1.0, 1.5), non-terminal deadlock, randomized backoff, long episodes (max_timesteps = 100), memory-with-communication.
- Random baseline for floor comparison (
experiments/models/random_baseline.py). - Statistics: 95% confidence intervals (Wald for proportions, normal-approximation for means) and Mann-Whitney U tests for cross-condition significance (
experiments/scripts/statistics.py). - Aggregated results in
experiments/results/aggregated/(master CSV, per-analysis tables, per-episode data). - Reproducibility scripts:
run.py,aggregate.py,generate_figures.pyunderexperiments/scripts/. - Test suite: 34 tests covering the environment, metrics, and multi-round communication (
pytest tests/).
Changed
- Unified provider access through a single OpenRouter interface (replaces per-provider modules).
- Prompt templates reorganised into named variant directories under
experiments/prompts/.
Removed
- Per-provider model adapters (OpenAI, Anthropic, Google, xAI), superseded by the OpenRouter interface.
Installation
pip install dpbench==0.2.0