Skip to content
#

agent-benchmarks

Here is 1 public repository matching this topic...

Language: JavaScript
Filter by language

用同一套提示词评测多模型 Agent 从零实现完整可玩游戏的能力:固定技术栈、固定评分标准、按模型归档对比。| A standardized prompt for benchmarking LLM agents on building a complete game from scratch — fixed tech stack, fixed rubric, per-model results.

  • Updated Aug 14, 2026
  • JavaScript

Improve this page

Add a description, image, and links to the agent-benchmarks topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the agent-benchmarks topic, visit your repo's landing page and select "manage topics."

Learn more