What’s new in MLPerf Client v2.0
MLPerf Client v2.0 expands the benchmark's scope to reflect the latest advancements in AI PC capabilities:
Image Generation
- New dedicated category for generative visual tasks, featuring the Flux 2 Klein 4B model (experimental).
Agentic AI Benchmarking
- New support for measuring the performance of AI agents in real-world scenarios, including Software Engineering (SWE) and Data Analyst tasks. Benchmarks track end-to-end performance, including LLM inference and tool execution time.
Enhanced LLM Benchmarking
- Updated model lineup: Llama 3.1 8B and Phi 4 Mini Instruct are now the mandatory base benchmarks (Phi 3.5 has been removed).
- Added Qwen 3 8B as an experimental test.
- Phi 4 Reasoning 14B has been moved to the extended category.
- Mandatory support for 4K prompt lengths in base benchmarks.
- Transitioned select summarization tasks to structured JSON output tasks.
Tooling & Performance Improvements
- Continued optimizations for cross-platform hardware acceleration on Windows, macOS, and Linux.
- Enhanced CLI capabilities, including CSV export for easier results analysis.
- Improved GUI responsiveness, stability, and memory utilization tracking.
Known issues
- No known issues at this time.