Skip to content

Benchmarks and Results

Sami Bajwa edited this page Jun 1, 2026 · 1 revision

Benchmarks and Results

Real measured token counts. All tests run with claude-sonnet-4-20250514 using the cl100k_base tokenizer.


Benchmark 1 — Prompt Compression Savings

Task Verbose Tokens Compressed Tokens Savings
Code debug request 53 8 85%
Blog post request 44 18 59%
Data extraction 49 19 61%
Email draft 58 22 62%
Code generation 41 14 66%
Meeting summary 46 16 65%
ELI5 explanation 35 12 66%
Social media content 51 20 61%
Comparison request 42 15 64%
Proofreading 37 10 73%

Average input token savings: 66%


Benchmark 2 — Output Format Control

Condition Avg Response Tokens Reduction
No format instruction (baseline) 487
Format: bullet points 312 36%
Bullet points + Anti-Filler Block 241 51%
max: 200 words constraint 198 59%
Format: JSON only 143 71%
Full format + length + negative combo 156 68%

Benchmark 3 — XML Structuring vs Unstructured

Prompt Style Clarification Rounds Completed First Try
Unstructured prose 2.3 avg 4 / 10
Structured headings 1.1 avg 7 / 10
XML tags 0.2 avg 9 / 10

Benchmark 4 — Context Management

Session Length Full History Tokens State Summary Tokens Savings
5 messages ~1,500 85 94%
10 messages ~3,500 92 97%
20 messages ~9,000 110 99%

Benchmark 5 — One-Shot Examples

Condition Format Match Rate Instruction Tokens
No example, verbose description 6 / 10 ~140
One-shot example 9 / 10 ~45

Overall Summary

Technique Input Savings Output Savings
Prompt Compression 59–85%
Output Format Control 36–71%
XML Structuring Eliminates 0–2 rounds
Context Pruning 94–99%
One-Shot Examples 68% instruction tokens

Combined effect: 3–8x fewer total tokens per session.

Clone this wiki locally