用同一套提示词评测多模型 Agent 从零实现完整可玩游戏的能力:固定技术栈、固定评分标准、按模型归档对比。| A standardized prompt for benchmarking LLM agents on building a complete game from scratch — fixed tech stack, fixed rubric, per-model results.
nodejs benchmark canvas game-development web-audio code-generation tank-battle llm-agents agent-benchmark agent-benchmarks
-
Updated
Aug 14, 2026 - JavaScript