Skip to content

Releases: WorldFlowAI/sembench

sembench 0.2.3

Choose a tag to compare

@zbennett10 zbennett10 released this 24 Sep 01:38

0.2.3

  • The system turn a manifest row asks for is now sent. Rows that carry
    metadata.system_prompt are replayed as [system, user], the message list
    their expectations were measured against. Earlier releases sent the user
    turn alone.
  • --staged-prefill replays a row's metadata.staged_prefill plan: each
    cut is sent as a token-exact /v1/completions prefix (max_tokens=1,
    vllm_xargs.semblend_stage=1), then the full prompt as token ids. TTFT
    counts from the first stage.
  • --capture-hint-role marks rows of one manifest role with
    vllm_xargs.semblend_capture=1. The harness knows the roles, so this
    measures the ceiling of hinted capture, not what a router could do.
  • Single-population alignment rate published beside the existing metrics
    (alignment_given_opportunity_in_class).
  • Test fix: the throughput check allows for wall_seconds rounding.

SemBench 0.2.0

Choose a tag to compare

@zbennett10 zbennett10 released this 03 Sep 20:08
  • run-load: donor->recipient streams with a fixed number in flight; requests per second, output tokens per second, TTFT percentiles for donors and recipients.
  • run-live-gateway --post-donor-delay-ms for engines that index donors asynchronously.
  • ROI cost model recalibrated to 8K to 24K measurements on an A10G.
  • Long-context manifest generator and the exact arm recipe in examples/overhead-benefit/.