Skip to content

Releases: postmanlabs/APIFlow-Bench

APIFlow-Bench 1.0

Choose a tag to compare

@zelinPostman zelinPostman released this 21 Jul 03:14

APIFlow-Bench 1.0 — first public release of the agentic API-workflow benchmark.

Release provenance (bank content sha256, epoch count, panel configs) is pinned in registry.json.

Transcripts — qwen3p7-plus

Choose a tag to compare

Raw per-trial transcripts for qwen3p7-plus on APIFlow-Bench 1.0.

  • Bulk download: transcripts-qwen3p7-plus.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/qwen3p7-plus/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.

Transcripts — qwen3p6-plus

Choose a tag to compare

Raw per-trial transcripts for qwen3p6-plus on APIFlow-Bench 1.0.

  • Bulk download: transcripts-qwen3p6-plus.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/qwen3p6-plus/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.

Transcripts — minimax-m2p7

Choose a tag to compare

Raw per-trial transcripts for minimax-m2p7 on APIFlow-Bench 1.0.

  • Bulk download: transcripts-minimax-m2p7.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/minimax-m2p7/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.

Transcripts — kimi-k2p7-code

Choose a tag to compare

Raw per-trial transcripts for kimi-k2p7-code on APIFlow-Bench 1.0.

  • Bulk download: transcripts-kimi-k2p7-code.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/kimi-k2p7-code/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.

Transcripts — kimi-k2p6

Choose a tag to compare

Raw per-trial transcripts for kimi-k2p6 on APIFlow-Bench 1.0.

  • Bulk download: transcripts-kimi-k2p6.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/kimi-k2p6/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.

Transcripts — kimi-k2p5

Choose a tag to compare

Raw per-trial transcripts for kimi-k2p5 on APIFlow-Bench 1.0.

  • Bulk download: transcripts-kimi-k2p5.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/kimi-k2p5/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.

Transcripts — gpt-oss-20b

Choose a tag to compare

Raw per-trial transcripts for gpt-oss-20b on APIFlow-Bench 1.0.

  • Bulk download: transcripts-gpt-oss-20b.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/gpt-oss-20b/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.

Transcripts — gpt-oss-120b

Choose a tag to compare

Raw per-trial transcripts for gpt-oss-120b on APIFlow-Bench 1.0.

  • Bulk download: transcripts-gpt-oss-120b.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/gpt-oss-120b/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.

Transcripts — gpt-5.6-terra

Choose a tag to compare

Raw per-trial transcripts for gpt-5.6-terra on APIFlow-Bench 1.0.

  • Bulk download: transcripts-gpt-5.6-terra.tar.gz (attached) — every trial for this model.
  • Individual files: served from the transcripts repo — https://github.com/postmanlabs/apiflow-bench-transcripts/raw/main/gpt-5.6-terra/<task_id>-t<N>.json (trial numbering matches the leaderboard's #trial-N).

Each JSON carries the full conversation (messages, reasoning blocks, tool calls and results), token usage, duration, and the verifier output as stored.