📄 Paper link: arXiv
The growth of agent skills has transformed how agentic systems are built, evaluated, and deployed. As skill libraries continue to scale, rigorous evaluation becomes critical to ensuring their utility, quality, and safety in real-world applications. Consequently, the field is undergoing an emerging paradigm shift from isolated skill creation to automated, evaluation-driven skill evolution. In this survey, we systematically examine the landscape of skill evolution and evaluation beyond foundational skill creation. We categorize evolution into four distinct paradigms, spanning execution feedback, trajectory distillation, compression, and reinforcement learning, showing how each element contributes to improving skill utility and reliability. We also provide an analysis of six skill-centric benchmark categories, identifying structural gaps in benchmark coverage, trade-offs, and metric richness to advance skill research. Finally, we identify open directions for building skill ecosystems that are generalizable, efficient, and verifiably safe.
This repository accompanies our survey on the lifecycle of agent skills for LLM agents — spanning skill creation & usage, evolution, and evaluation. Below, we provide an overview figure and curated reference tables for the works discussed in the survey.
We map the landscape of agent skill evolution strategies (§3) through comparative analysis and design recommendations. We offer evaluation insights (§4) through structural gaps and benchmark limitations, and open challenges (§5) for robust real-world skill deployment.
- 🛠️ Skill Creation, Usage & Evolution
- 📊 Skill-Centric Evaluation & Benchmarks
- 🌐 General-Domain Evaluation Benchmarks
- 📝 Citation
This table covers the works discussed in the skill definition & creation, skill usage (retrieval, routing, management), and skill evolution sections.
| Work | Reference | Category | Year |
|---|---|---|---|
| Anthropic Skill Creator | Link | 🤖 Creation — Automatic Creation | 2025 |
| Voyager | Link | 🤖 Creation — Automatic Creation | 2023 |
| SAGE | Link | 🎯 Creation — RL-based Creation | 2025 |
| ARISE | Link | 🎯 Creation — RL-based Creation | 2026 |
| SkillRouter | Link | 🔍 Usage — Retrieval / Routing | 2026 |
| SkillFlow | Link | 🔍 Usage — Retrieval | 2025 |
| SkillOrchestra | Link | 🔍 Usage — Routing | 2026 |
| AgentSkillOS | Link | 🗂️ Usage — Management / Orchestration | 2026 |
| SSL | Link | 🗂️ Usage — Management | 2026 |
| SkillForge | Link | 🔄 Evolution — Execution Feedback | 2026 |
| CoEvoSkills | Link | 🔄 Evolution — Execution Feedback | 2026 |
| EmbodiSkill | Link | 🔄 Evolution — Execution Feedback | 2026 |
| Skills-Coach | Link | 🔄 Evolution — Execution Feedback | 2026 |
| Ctx2Skill | Link | 🔄 Evolution — Execution Feedback | 2026 |
| AutoSkill | Link | 🔄 Evolution — Execution Feedback | 2026 |
| SkillClaw | Link | 🔄 Evolution — Execution Feedback | 2026 |
| SPARK | Link | 🧪 Evolution — Trajectory Distillation | 2026 |
| Trace2Skill | Link | 🧪 Evolution — Trajectory Distillation | 2026 |
| Memento-Skills | Link | 🧪 Evolution — Trajectory Distillation | 2026 |
| XSkill | Link | 🧪 Evolution — Trajectory Distillation | 2026 |
| SkillNet | Link | 🗜️ Evolution — Compression & Augmentation | 2026 |
| SkillX | Link | 🗜️ Evolution — Compression & Augmentation | 2026 |
| SkillFoundry | Link | 🗜️ Evolution — Compression & Augmentation | 2026 |
| SkillReducer | Link | 🗜️ Evolution — Compression & Augmentation | 2026 |
| D2Skill | Link | 🏋️ Evolution — Reinforcement Learning | 2026 |
| SkillRL | Link | 🏋️ Evolution — Reinforcement Learning | 2026 |
| Skill1 | Link | 🏋️ Evolution — Reinforcement Learning | 2026 |
This table covers the benchmarks discussed in the skill-centric evaluation and benchmarks section.
| Work | Reference | Category | Year |
|---|---|---|---|
| SkillsBench | Link | ⚡ Skill Utility | 2026 |
| SkillCraft | Link | ⚡ Skill Utility | 2026 |
| SkillLearnBench | Link | ✨ Skill Generation | 2026 |
| SkillRouter | Link | 🔍 Retrieval & Routing | 2026 |
| SRA-Bench | Link | 🔍 Retrieval & Routing | 2026 |
| AgentSkillOS Benchmark | Link | 🔍 Retrieval & Routing | 2026 |
| SkillTester | Link | 🛡️ Safety & Security | 2026 |
| SkillGuardBench | Link | 🛡️ Safety & Security | 2026 |
| SKILL-INJECT | Link | 🛡️ Safety & Security | 2026 |
| SWE-Skills-Bench | Link | 💻 Software Engineering | 2026 |
| WildClawBench | Link | 🌍 Real-world Environment | 2026 |
| SkillForge | Link | 🌍 Real-world Environment | 2026 |
General-domain benchmarks that were not designed for skill evaluation but are still applicable to assessing agent skills, grouped by task category.
| Work | Reference | Category | Year |
|---|---|---|---|
| ALFWorld | Link | 🕹️ Interactive Agent Environments | 2020 |
| WebShop | Link | 🕹️ Interactive Agent Environments | 2022 |
| ScienceWorld | Link | 🕹️ Interactive Agent Environments | 2022 |
| WebArena | Link | 🕹️ Interactive Agent Environments | 2024 |
| AgentBench | Link | 🕹️ Interactive Agent Environments | 2024 |
| AppWorld | Link | 🕹️ Interactive Agent Environments | 2024 |
| Terminal-Bench | Link | 💻 Code Generation & Software Engineering | 2026 |
| HumanEval | Link | 💻 Code Generation & Software Engineering | 2021 |
| MBPP | Link | 💻 Code Generation & Software Engineering | 2021 |
| EffiBench-X | Link | 💻 Code Generation & Software Engineering | 2026 |
| PIE dataset | Link | 💻 Code Generation & Software Engineering | 2023 |
| AMC / AIME | Link | 🔢 Mathematical Reasoning | 2026 |
| Omni-MATH | Link | 🔢 Mathematical Reasoning | 2025 |
| DeepScaleR | Link | 🔢 Mathematical Reasoning | 2025 |
| Natural Questions (NQ) | Link | 💡 Question Answering & Knowledge-Intensive | 2019 |
| TriviaQA | Link | 💡 Question Answering & Knowledge-Intensive | 2017 |
| PopQA | Link | 💡 Question Answering & Knowledge-Intensive | 2023 |
| HotpotQA | Link | 💡 Question Answering & Knowledge-Intensive | 2018 |
| 2WikiMultiHopQA | Link | 💡 Question Answering & Knowledge-Intensive | 2020 |
| MuSiQue | Link | 💡 Question Answering & Knowledge-Intensive | 2022 |
| Bamboogle | Link | 💡 Question Answering & Knowledge-Intensive | 2023 |
| WikiTableQuestions | Link | 💡 Question Answering & Knowledge-Intensive | 2015 |
| GAIA | Link | 💡 Question Answering & Knowledge-Intensive | 2024 |
| HLE (Humanity's Last Exam) | Link | 💡 Question Answering & Knowledge-Intensive | 2025 |
| MMLU | Link | 📚 Knowledge, Language & Instruction-Following | 2020 |
| AlpacaEval | Link | 📚 Knowledge, Language & Instruction-Following | 2023 |
| MT-Bench | Link | 📚 Knowledge, Language & Instruction-Following | 2023 |
| WildBench | Link | 📚 Knowledge, Language & Instruction-Following | 2025 |
| LoCoMo | Link | 💬 Memory & Conversational | 2024 |
| LongMemEval | Link | 💬 Memory & Conversational | 2024 |
| StuLife | Link | 💬 Memory & Conversational | 2025 |
| VisualToolBench | Link | 🖼️ Multimodal & Tool-Use | 2025 |
| TIR-Bench | Link | 🖼️ Multimodal & Tool-Use | 2025 |
| MMSearch-Plus | Link | 🖼️ Multimodal & Tool-Use | 2025 |
| MMBrowseComp | Link | 🖼️ Multimodal & Tool-Use | 2025 |
| AgentVista | Link | 🖼️ Multimodal & Tool-Use | 2026 |
| SpreadsheetBench | Link | 🖼️ Multimodal & Tool-Use | 2024 |
| BFCL-v3 | Link | 🖼️ Multimodal & Tool-Use | 2024 |
| τ²-Bench | Link | 🖼️ Multimodal & Tool-Use | 2025 |
| M³-Bench | Link | 🖼️ Multimodal & Tool-Use | 2025 |
| MineDojo / Minecraft | Link | 🤸 Embodied / Open-Ended | 2022 |
@misc{ding2026skillsurvey,
title = {Agent Skill Evaluation and Evolution: Frameworks and Benchmarks},
author = {Ding, Kexin and Zhou, Yang and Jin, Can and Tong, Feng and
Zhou, Mu and Metaxas, Dimitris N.},
year = {2026},
eprint = {2606.11435},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2606.11435}
}
