Skip to content

Releases: najmulhasan-code/dpbench

v0.2.1

Choose a tag to compare

@najmulhasan-code najmulhasan-code released this 05 Jun 23:25

Documentation patch. No code changes.

Added

  • arXiv badge linking to the paper.
  • Citation entry with the official arXiv BibTeX in the README.
  • Downloads badge from pepy.tech.
  • "Read the full paper on arXiv" link in the Overview section.

Changed

  • The teaser figure is now clickable and links to the arXiv paper.
  • Removed the Python version badge (redundant with pyproject.toml's requires-python).

Links

v0.2.0

Choose a tag to compare

@najmulhasan-code najmulhasan-code released this 03 Jun 17:40

Feature release. Expands the v0.1.0 evaluation (4 models, 8 conditions) into a layered suite with independent control of action protocol, communication structure, group size, and prompting strategy.

Added

  • Three-layer experiment suite (experiments/configs/conditions.yaml):
    Layer 1 universality (six baseline conditions: simultaneous/sequential × communication/no-communication × N ∈ {5, 10}, run across all models),
    Layer 2 scaling (extends group size to N ∈ {7, 15, 20}),
    Layer 3 ablations.
  • Ablations: prompt variant (symmetry breaking, theory of mind, resource ordering, minimal), memory window (k ∈ {3, 5}), multi-round communication (3, 5 rounds), temperature (0.3, 1.0, 1.5), non-terminal deadlock, randomized backoff, long episodes (max_timesteps = 100), memory-with-communication.
  • Random baseline for floor comparison (experiments/models/random_baseline.py).
  • Statistics: 95% confidence intervals (Wald for proportions, normal-approximation for means) and Mann-Whitney U tests for cross-condition significance (experiments/scripts/statistics.py).
  • Aggregated results in experiments/results/aggregated/ (master CSV, per-analysis tables, per-episode data).
  • Reproducibility scripts: run.py, aggregate.py, generate_figures.py under experiments/scripts/.
  • Test suite: 34 tests covering the environment, metrics, and multi-round communication (pytest tests/).

Changed

  • Unified provider access through a single OpenRouter interface (replaces per-provider modules).
  • Prompt templates reorganised into named variant directories under experiments/prompts/.

Removed

  • Per-provider model adapters (OpenAI, Anthropic, Google, xAI), superseded by the OpenRouter interface.

Installation

pip install dpbench==0.2.0

v0.1.0 - Initial Release

Choose a tag to compare

@najmulhasan-code najmulhasan-code released this 29 Mar 14:53

Initial release of DPBench - A benchmark framework for evaluating LLM multi-agent coordination under resource contention.

Features

  • Simultaneous and sequential decision modes
  • Configurable agent counts (N philosophers)
  • Optional inter-agent communication
  • Comprehensive metrics: deadlock rate, throughput, fairness, starvation
  • JSONL logging and human-readable transcripts
  • Works with any LLM via simple function protocol