Skip to content

v0.2.0

Choose a tag to compare

@najmulhasan-code najmulhasan-code released this 03 Jun 17:40
· 3 commits to main since this release

Feature release. Expands the v0.1.0 evaluation (4 models, 8 conditions) into a layered suite with independent control of action protocol, communication structure, group size, and prompting strategy.

Added

  • Three-layer experiment suite (experiments/configs/conditions.yaml):
    Layer 1 universality (six baseline conditions: simultaneous/sequential × communication/no-communication × N ∈ {5, 10}, run across all models),
    Layer 2 scaling (extends group size to N ∈ {7, 15, 20}),
    Layer 3 ablations.
  • Ablations: prompt variant (symmetry breaking, theory of mind, resource ordering, minimal), memory window (k ∈ {3, 5}), multi-round communication (3, 5 rounds), temperature (0.3, 1.0, 1.5), non-terminal deadlock, randomized backoff, long episodes (max_timesteps = 100), memory-with-communication.
  • Random baseline for floor comparison (experiments/models/random_baseline.py).
  • Statistics: 95% confidence intervals (Wald for proportions, normal-approximation for means) and Mann-Whitney U tests for cross-condition significance (experiments/scripts/statistics.py).
  • Aggregated results in experiments/results/aggregated/ (master CSV, per-analysis tables, per-episode data).
  • Reproducibility scripts: run.py, aggregate.py, generate_figures.py under experiments/scripts/.
  • Test suite: 34 tests covering the environment, metrics, and multi-round communication (pytest tests/).

Changed

  • Unified provider access through a single OpenRouter interface (replaces per-provider modules).
  • Prompt templates reorganised into named variant directories under experiments/prompts/.

Removed

  • Per-provider model adapters (OpenAI, Anthropic, Google, xAI), superseded by the OpenRouter interface.

Installation

pip install dpbench==0.2.0