Repository navigation
What's New
- Server benchmarks (RTX 4000 SFF Ada 20GB)
- Multi-language report generation (EN, DE, FR, ES, FA, AR)
- Prioritized server benchmark charts in README
- Fixed npm/GitHub Packages publish workflow
Server Benchmark Results
| Model | Throughput | TTFT |
|---|---|---|
| GPT-2 | 407.1 tok/s | 4.0ms |
| Qwen2.5-0.5B | 140.7 tok/s | 10.9ms |
| TinyLlama-1.1B | 93.0 tok/s | 30.6ms |
| Phi-1.5 | 78.8 tok/s | 37.2ms |