Releases: najmulhasan-code/dpbench
Releases · najmulhasan-code/dpbench
Release list
v0.2.1
Documentation patch. No code changes.
Added
- arXiv badge linking to the paper.
- Citation entry with the official arXiv BibTeX in the README.
- Downloads badge from
pepy.tech. - "Read the full paper on arXiv" link in the Overview section.
Changed
- The teaser figure is now clickable and links to the arXiv paper.
- Removed the Python version badge (redundant with
pyproject.toml'srequires-python).
Links
v0.2.0
Feature release. Expands the v0.1.0 evaluation (4 models, 8 conditions) into a layered suite with independent control of action protocol, communication structure, group size, and prompting strategy.
Added
- Three-layer experiment suite (
experiments/configs/conditions.yaml):
Layer 1 universality (six baseline conditions: simultaneous/sequential × communication/no-communication × N ∈ {5, 10}, run across all models),
Layer 2 scaling (extends group size to N ∈ {7, 15, 20}),
Layer 3 ablations. - Ablations: prompt variant (symmetry breaking, theory of mind, resource ordering, minimal), memory window (k ∈ {3, 5}), multi-round communication (3, 5 rounds), temperature (0.3, 1.0, 1.5), non-terminal deadlock, randomized backoff, long episodes (max_timesteps = 100), memory-with-communication.
- Random baseline for floor comparison (
experiments/models/random_baseline.py). - Statistics: 95% confidence intervals (Wald for proportions, normal-approximation for means) and Mann-Whitney U tests for cross-condition significance (
experiments/scripts/statistics.py). - Aggregated results in
experiments/results/aggregated/(master CSV, per-analysis tables, per-episode data). - Reproducibility scripts:
run.py,aggregate.py,generate_figures.pyunderexperiments/scripts/. - Test suite: 34 tests covering the environment, metrics, and multi-round communication (
pytest tests/).
Changed
- Unified provider access through a single OpenRouter interface (replaces per-provider modules).
- Prompt templates reorganised into named variant directories under
experiments/prompts/.
Removed
- Per-provider model adapters (OpenAI, Anthropic, Google, xAI), superseded by the OpenRouter interface.
Installation
pip install dpbench==0.2.0
v0.1.0 - Initial Release
Initial release of DPBench - A benchmark framework for evaluating LLM multi-agent coordination under resource contention.
Features
- Simultaneous and sequential decision modes
- Configurable agent counts (N philosophers)
- Optional inter-agent communication
- Comprehensive metrics: deadlock rate, throughput, fairness, starvation
- JSONL logging and human-readable transcripts
- Works with any LLM via simple function protocol