Skip to content

Latest commit

 

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

🤖 Agent Skill Evaluation and Evolution: Frameworks and Benchmarks

📄 Paper link: arXiv

The growth of agent skills has transformed how agentic systems are built, evaluated, and deployed. As skill libraries continue to scale, rigorous evaluation becomes critical to ensuring their utility, quality, and safety in real-world applications. Consequently, the field is undergoing an emerging paradigm shift from isolated skill creation to automated, evaluation-driven skill evolution. In this survey, we systematically examine the landscape of skill evolution and evaluation beyond foundational skill creation. We categorize evolution into four distinct paradigms, spanning execution feedback, trajectory distillation, compression, and reinforcement learning, showing how each element contributes to improving skill utility and reliability. We also provide an analysis of six skill-centric benchmark categories, identifying structural gaps in benchmark coverage, trade-offs, and metric richness to advance skill research. Finally, we identify open directions for building skill ecosystems that are generalizable, efficient, and verifiably safe.

This repository accompanies our survey on the lifecycle of agent skills for LLM agents — spanning skill creation & usage, evolution, and evaluation. Below, we provide an overview figure and curated reference tables for the works discussed in the survey.

Overview of agent skill creation/usage, evolution, and evaluation

We map the landscape of agent skill evolution strategies (§3) through comparative analysis and design recommendations. We offer evaluation insights (§4) through structural gaps and benchmark limitations, and open challenges (§5) for robust real-world skill deployment.

📋 Table of Contents


🛠️ Skill Creation, Usage & Evolution

This table covers the works discussed in the skill definition & creation, skill usage (retrieval, routing, management), and skill evolution sections.

Work Reference Category Year
Anthropic Skill Creator Link 🤖 Creation — Automatic Creation 2025
Voyager Link 🤖 Creation — Automatic Creation 2023
SAGE Link 🎯 Creation — RL-based Creation 2025
ARISE Link 🎯 Creation — RL-based Creation 2026
SkillRouter Link 🔍 Usage — Retrieval / Routing 2026
SkillFlow Link 🔍 Usage — Retrieval 2025
SkillOrchestra Link 🔍 Usage — Routing 2026
AgentSkillOS Link 🗂️ Usage — Management / Orchestration 2026
SSL Link 🗂️ Usage — Management 2026
SkillForge Link 🔄 Evolution — Execution Feedback 2026
CoEvoSkills Link 🔄 Evolution — Execution Feedback 2026
EmbodiSkill Link 🔄 Evolution — Execution Feedback 2026
Skills-Coach Link 🔄 Evolution — Execution Feedback 2026
Ctx2Skill Link 🔄 Evolution — Execution Feedback 2026
AutoSkill Link 🔄 Evolution — Execution Feedback 2026
SkillClaw Link 🔄 Evolution — Execution Feedback 2026
SPARK Link 🧪 Evolution — Trajectory Distillation 2026
Trace2Skill Link 🧪 Evolution — Trajectory Distillation 2026
Memento-Skills Link 🧪 Evolution — Trajectory Distillation 2026
XSkill Link 🧪 Evolution — Trajectory Distillation 2026
SkillNet Link 🗜️ Evolution — Compression & Augmentation 2026
SkillX Link 🗜️ Evolution — Compression & Augmentation 2026
SkillFoundry Link 🗜️ Evolution — Compression & Augmentation 2026
SkillReducer Link 🗜️ Evolution — Compression & Augmentation 2026
D2Skill Link 🏋️ Evolution — Reinforcement Learning 2026
SkillRL Link 🏋️ Evolution — Reinforcement Learning 2026
Skill1 Link 🏋️ Evolution — Reinforcement Learning 2026

📊 Skill-Centric Evaluation & Benchmarks

This table covers the benchmarks discussed in the skill-centric evaluation and benchmarks section.

Work Reference Category Year
SkillsBench Link ⚡ Skill Utility 2026
SkillCraft Link ⚡ Skill Utility 2026
SkillLearnBench Link ✨ Skill Generation 2026
SkillRouter Link 🔍 Retrieval & Routing 2026
SRA-Bench Link 🔍 Retrieval & Routing 2026
AgentSkillOS Benchmark Link 🔍 Retrieval & Routing 2026
SkillTester Link 🛡️ Safety & Security 2026
SkillGuardBench Link 🛡️ Safety & Security 2026
SKILL-INJECT Link 🛡️ Safety & Security 2026
SWE-Skills-Bench Link 💻 Software Engineering 2026
WildClawBench Link 🌍 Real-world Environment 2026
SkillForge Link 🌍 Real-world Environment 2026

🌐 Table 3 — General-Domain Evaluation Benchmarks

General-domain benchmarks that were not designed for skill evaluation but are still applicable to assessing agent skills, grouped by task category.

Work Reference Category Year
ALFWorld Link 🕹️ Interactive Agent Environments 2020
WebShop Link 🕹️ Interactive Agent Environments 2022
ScienceWorld Link 🕹️ Interactive Agent Environments 2022
WebArena Link 🕹️ Interactive Agent Environments 2024
AgentBench Link 🕹️ Interactive Agent Environments 2024
AppWorld Link 🕹️ Interactive Agent Environments 2024
Terminal-Bench Link 💻 Code Generation & Software Engineering 2026
HumanEval Link 💻 Code Generation & Software Engineering 2021
MBPP Link 💻 Code Generation & Software Engineering 2021
EffiBench-X Link 💻 Code Generation & Software Engineering 2026
PIE dataset Link 💻 Code Generation & Software Engineering 2023
AMC / AIME Link 🔢 Mathematical Reasoning 2026
Omni-MATH Link 🔢 Mathematical Reasoning 2025
DeepScaleR Link 🔢 Mathematical Reasoning 2025
Natural Questions (NQ) Link 💡 Question Answering & Knowledge-Intensive 2019
TriviaQA Link 💡 Question Answering & Knowledge-Intensive 2017
PopQA Link 💡 Question Answering & Knowledge-Intensive 2023
HotpotQA Link 💡 Question Answering & Knowledge-Intensive 2018
2WikiMultiHopQA Link 💡 Question Answering & Knowledge-Intensive 2020
MuSiQue Link 💡 Question Answering & Knowledge-Intensive 2022
Bamboogle Link 💡 Question Answering & Knowledge-Intensive 2023
WikiTableQuestions Link 💡 Question Answering & Knowledge-Intensive 2015
GAIA Link 💡 Question Answering & Knowledge-Intensive 2024
HLE (Humanity's Last Exam) Link 💡 Question Answering & Knowledge-Intensive 2025
MMLU Link 📚 Knowledge, Language & Instruction-Following 2020
AlpacaEval Link 📚 Knowledge, Language & Instruction-Following 2023
MT-Bench Link 📚 Knowledge, Language & Instruction-Following 2023
WildBench Link 📚 Knowledge, Language & Instruction-Following 2025
LoCoMo Link 💬 Memory & Conversational 2024
LongMemEval Link 💬 Memory & Conversational 2024
StuLife Link 💬 Memory & Conversational 2025
VisualToolBench Link 🖼️ Multimodal & Tool-Use 2025
TIR-Bench Link 🖼️ Multimodal & Tool-Use 2025
MMSearch-Plus Link 🖼️ Multimodal & Tool-Use 2025
MMBrowseComp Link 🖼️ Multimodal & Tool-Use 2025
AgentVista Link 🖼️ Multimodal & Tool-Use 2026
SpreadsheetBench Link 🖼️ Multimodal & Tool-Use 2024
BFCL-v3 Link 🖼️ Multimodal & Tool-Use 2024
τ²-Bench Link 🖼️ Multimodal & Tool-Use 2025
M³-Bench Link 🖼️ Multimodal & Tool-Use 2025
MineDojo / Minecraft Link 🤸 Embodied / Open-Ended 2022

📝 Citation

@misc{ding2026skillsurvey,
  title         = {Agent Skill Evaluation and Evolution: Frameworks and Benchmarks},
  author        = {Ding, Kexin and Zhou, Yang and Jin, Can and Tong, Feng and
                   Zhou, Mu and Metaxas, Dimitris N.},
  year          = {2026},
  eprint        = {2606.11435},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2606.11435}
}

About

Agent Skill Evaluation and Evolution: Frameworks and Benchmarks

Resources

Stars

28 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors