You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Self-Improving Agents in the Era of Experience: A Survey of Self- to Meta-Evolution
We welcome issues and pull requests for missing work on harness agents, agent skills, memory, self-improvement, agent RL, evaluation, and safety.
News
[2026-06-25] 🎉 Release: Our survey is now available on OpenReview.
Citation
If you find this survey or paper list helpful, please cite our work:
@article{jiang2026selfimprovingagents,
title={Self-Improving Agents in the Era of Experience: A Survey of Self- to Meta-Evolution},
author={Che Jiang and Jincheng Zhong and Yu Fu and Kai Tian and Junlin Yang and Kaikai Zhao and Yuchong Wang and Tianwei Luo and Weizhi Wang and Yuxin Zuo and Guoli Jia and Xingtai Lv and Dianqiao Lei and Sihang Zeng and Yuru Wang and Zhenzhao Yuan and Xinwei Long and Ermo Hua and Can Ren and Xin Jiang and Shulei Xie and Yuanchun Zheng and Youbang Sun and Biqing Qi and Ning Ding and Kaiyan Zhang and Bowen Zhou},
journal={OpenReview Archive},
year={2026},
url={https://openreview.net/pdf?id=IUltZSgLMm}
}
This repository collects papers, systems, benchmarks, and resources for studying how deployed agentic AI systems become more capable after deployment.
We organize the landscape around the harness agent: a deployed runtime system whose behavior is jointly shaped by a base model, a mutable harness, a user-facing interface, and an environment-facing interface. Under this view, self-improvement first appears as fast runtime adaptation over external surfaces such as skills, memory, context, tools, and execution environments. Repeated experience may later be consolidated into model parameters through reinforcement learning, fine-tuning, or continual learning.
We organize the survey into four parts:
Paradigm shift: from task-bounded tool-use loops to deployed harness-centered runtime systems.
External path: skills, memory, context, tools, and environments as editable runtime adaptation surfaces.
Parameter path and meta-evolution: agent RL, continual learning, and meta-agents for durable learning and update orchestration.
Conditions and limits: evaluation, safety, governance, and open problems for reliable post-deployment improvement.
Paper List
This paper list follows the references cited by the LaTeX manuscript. It currently includes 331 unique cited entries from 379 unique manuscript citation keys and 379 cited BibTeX records.
Every row includes a date, display name, title, and at least one public source badge. Cited entries whose public source URL still needs verification are omitted from this table until complete metadata is available (44 currently omitted; 4 duplicate cited records collapsed).
Foundations and Surveys
Date
Name
Title
Paper
Github
2026-06
OpenSkill
OpenSkill: Open-World Self-Evolution for LLM Agents
-
2026-05
huang2026rawexperience
From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
-
2026-05
SkillOpt
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
-
2026-04
cursor2026cursor3
Meet the New Cursor
-
2026-04
neuralcomputers2026_paper
Neural Computers
-
2026-04
SWE-chat
SWE-chat: Coding Agent Interactions From Real Users in the Wild
-
2026-01
AI Agent Systems
AI Agent Systems: Architectures, Applications, and Evaluation
-
2025-08
fang2025selfevolvingagents
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
-
2025-07
A Survey of Self-Evolving Agents
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
-
2025-06
chen2025compoundaisystems
From Standalone LLMs to Integrated Intelligence: A Survey of Compound AI Systems
-
2025-02
anthropic2025claudecode
Claude 3.7 Sonnet and Claude Code
-
2025-01
silver2025eraexperience
Welcome to the Era of Experience
-
2024-05
WildChat
WildChat: 1M ChatGPT Interaction Logs in the Wild
-
2024-04
tao2024selfevolutionsurvey
A Survey on Self-Evolution of Large Language Models
-
2024-01
langchain2024langgraph
LangGraph
-
2023-09
xi2023riseagentsurvey
The Rise and Potential of Large Language Model Based Agents: A Survey
-
2023-08
wang2023llmagentsurvey
A Survey on Large Language Model based Autonomous Agents
-
2022-01
ahn2022can
Do as i can, not as i say: Grounding language in robotic affordances
-
Harness and Runtime Architecture
Date
Name
Title
Paper
Github
2026-08
Ouroboros
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
2026-06
recursive2026automatedresearch
First Steps Toward Automated AI Research
-
2026-06
osmani2026loopengineering
Loop Engineering
-
2026-06
Traj-Evolve
Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection
-
2026-05
Agent Harness Engineering
Agent Harness Engineering: A Survey
-
2026-04
meng2026agentharness
Agent Harness for Large Language Model Agents: A Survey
-
2026-04
boeckeler2026harnessengineering
Harness Engineering for Coding Agent Users
-
2026-04
martin2026harnessfailures
Most AI Agent Failures Are Harness Failures
-
2026-04
Scaling Managed Agents
Scaling Managed Agents: Decoupling the Brain from the Hands
-
2026-04
xu2026futureagentsopensource
The Future of Agents is Open Source (part 1 of 2)
-
2026-03
cursor2026composer2
Composer 2 Technical Report
-
2026-03
yue2026workflowsurvey
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
-
2026-02
Unlocking the Codex harness
Unlocking the Codex harness: how we built the App Server
-
2026-01
openharness2026_repo
OpenHarness
2025-11
young2025effectiveharnesses
Effective Harnesses for Long-Running Agents
-
2025-07
mei2025contextengineering
A Survey of Context Engineering for Large Language Models
-
2025-05
Darwin Godel Machine
Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
-
2025-05
openai2025codex
Introducing Codex
-
2020-01
Artificial Intelligence
Artificial Intelligence: A Modern Approach
-
2003-01
Goedel Machines
Goedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements
-
Skills and Skill Libraries
Date
Name
Title
Paper
Github
2026-07
Skill-SP
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
2026-06
li2026agenticenvironmentengineering
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
Group of Skills: Group-Structured Skill Retrieval for Agent Skill Libraries
-
2026-05
MIND-Skill
MIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction
-
2026-05
MUSE-Autoskill
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
-
2026-05
OpenClaw Research
OpenClaw Research: A Systematic Survey of Large Language Model Agents in Open Deployment
-
2026-05
SkillEvolver
SkillEvolver: Skill Learning as a Meta-Skill
-
2026-05
SkillRAE
SkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution
-
2026-05
SkillRet
SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
-
2026-04
CoEvoSkills
CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
-
2026-04
Corpus2Skill
Corpus2Skill: Distilling Document Corpora into Hierarchical Skill Directories for Agent Navigation
-
2026-04
Externalization in LLM Agents
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
-
2026-04
hivemind2026_repo
Hivemind: Continual Learning Layer that Distills Coding-Agent Session Trajectories into Reusable Skills
2026-04
skillswild2026realistic
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
-
2026-04
SKILL0
SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
-
2026-04
SkillX
SkillX: Automatically Constructing Skill Knowledge Bases for Agents
-
2026-03
bi2026repositorymining
Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction
-
2026-03
AutoSkill
AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
-
2026-03
d2skill2026dynamic
Dynamic Dual-Granularity Skill Bank for Agentic RL
-
2026-03
EvoSkill
EvoSkill: Automated Skill Discovery for Multi-Agent Systems
-
2026-03
From Model to Agent
From Model to Agent: Equipping the Responses API with a Computer Environment
-
2026-03
MetaClaw
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
-
2026-03
SkillNet
SkillNet: Create, Evaluate, and Connect AI Skills
-
2026-03
SkillRouter
SkillRouter: Retrieve-and-Rerank Skill Selection for LLM Agents at Scale
-
2026-03
SWE-Skills-Bench
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
-
2026-03
Trace2Skill
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
-
2026-03
XSkill
XSkill: Continual Learning from Experience and Skills in Multimodal Agents
-
2026-02
Skill-Pro
Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
-
2026-02
SkillRL
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
-
2026-02
SkillsBench
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
-
2026-01
AutoRefine
AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement
-
2026-01
hermesagent2026_repo
Hermes Agent
2026-01
li2026singleagentskills
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
-
2025-12
sage2025selfimproving
Reinforcement Learning for Self-Improving Agent with Skill Library
-
2025-10
composeincontext2025skills
Can Language Models Compose Skills In-Context?
-
Memory and Context Management
Date
Name
Title
Paper
Github
2026-05
Auto-Dreamer
Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents
-
2026-05
MemORAI
MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents
-
2026-05
MEMOREPAIR
MEMOREPAIR: Barrier-First Cascade Repair in Agentic Memory
-
2026-05
zou2026demem
Remember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory
-
2026-05
SAGE
SAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory
-
2026-04
APEX-MEM
APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI
-
2026-04
Cognis
Cognis: Context-Aware Memory for Conversational AI Agents
-
2026-04
DeltaMem
DeltaMem: Towards Agentic Memory Management via Reinforcement Learning
-
2026-04
zhang2026lightmem
Lightweight LLM Agent Memory with Small Language Models
-
2026-03
zhang2026amac
Adaptive Memory Admission Control for LLM Agents
-
2026-03
AriadneMem
AriadneMem: Threading the Maze of Lifelong Memory for LLM Agents
-
2026-03
Chronos
Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory
-
2026-03
CLAG
CLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents
-
2026-03
GAAMA
GAAMA: Graph Augmented Associative Memory for Agents
-
2026-03
Memori
Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents
-
2026-03
Memory for Autonomous LLM Agents
Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers
-
2026-03
PlugMem
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
-
2026-02
CAST
CAST: Character-and-Scene Episodic Memory for Agents
-
2026-02
From Lossy to Verified
From Lossy to Verified: A Provenance-Aware Tiered Memory for Agents
-
2026-02
Live-Evo
Live-Evo: Online Evolution of Agentic Memory from Continuous Feedback
-
2026-02
MemoryArena
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
-
2026-02
xinlewu2026umem
Towards Autonomous Memory Agents
-
2026-02
UI-Mem
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents
-
2026-02
xMemory
xMemory: Beyond RAG for Agent Memory -- Retrieval by Decoupling and Aggregation
-
2026-01
Active Context Compression
Active Context Compression: Autonomous Memory Management in LLM Agents
-
2026-01
Agentic Memory
Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
-
2026-01
EMemBench
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
-
2026-01
H-Mem
H-Mem: A Novel Memory Mechanism for Evolving and Retrieving Agent Memory via a Hybrid Structure
-
2026-01
baldelli2026hangman
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
-
2026-01
Mem2ActBench
Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
-
2026-01
MemRL
MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
-
2026-01
SimpleMem
SimpleMem: Efficient Lifelong Memory for LLM Agents
-
2025-11
WebCoach
WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
-
2025-08
Nemori
Nemori: Self-Organizing Agent Memory Inspired by Cognitive Science
-
2025-03
Search-R1
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
-
2024-10
LongMemEval
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
-
2024-04
zhang2024memorymechanism
A Survey on the Memory Mechanism of Large Language Model Based Agents
-
2023-12
empoweringworkingmemory2023
Empowering Working Memory for Large Language Model Agents
-
2023-10
MemGPT
MemGPT: Towards LLMs as Operating Systems
-
2023-03
Reflexion
Reflexion: Language Agents with Verbal Reinforcement Learning
-
2020-05
lewis2020rag
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
-
2016-12
kirkpatrick2017ewc
Overcoming catastrophic forgetting in neural networks
-
Environments, Tools, and Runtime Feedback
Date
Name
Title
Paper
Github
2026-05
zhong2026executablebenchmark
An Executable Benchmarking Suite for Tool-Using Agents
-
2026-05
wu2026chemcost
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
-
2026-05
CUA-Gym
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
-
2026-05
MANTRA
MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents
-
2026-05
PhysicianBench
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
-
2026-05
SkillSmith
SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces
-
2026-05
When Simulation Lies
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
-
2026-04
Agent-World
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
-
2026-04
Agentic World Modeling
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
-
2026-04
Gym-Anything
Gym-Anything: Turn any Software into an Agent Environment
-
2026-04
zhou2026sandmle
Synthetic Sandbox for Training Machine Learning Engineering Agents
-
2026-04
ToolMisuseBench
ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems
-
2026-02
CLI-Gym
CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
-
2026-02
shen2026wac
World-Model-Augmented Web Agents with Action Correction
-
2026-01
xiang2026selfevolvingcoevolution
A Systematic Survey of Self-Evolving Agents: From Model-Centric to Environment-Driven Co-Evolution
-
2026-01
AG-UI
AG-UI: The Agent-User Interaction Protocol
2026-01
a2a_spec_2026
Agent2Agent (A2A) Protocol
2026-01
CLI-Anything
CLI-Anything: Making ALL Software Agent-Native
2026-01
Harbor
Harbor: A Framework for Running Agent Evaluations and Creating and Using RL Environments
2026-01
LiteCoder-Terminal
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
-
2026-01
MEnvAgent
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering
-
2026-01
openclaw2026_repo
OpenClaw
2026-01
SearchGym
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
-
2026-01
Terminal-Bench
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
-
2026-01
WebGym
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
-
2025-08
SEAgent
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
-
2025-06
proceduraltooluse2025_paper
Procedural Environment Generation for Tool-Use Agents
-
2025-05
DeepResearchGym
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
-
2025-05
MLE-Dojo
MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
-
2025-03
chezelles2025browsergym
The BrowserGym Ecosystem for Web Agent Research
-
2025-01
mcp_spec_2025
Model Context Protocol Specification
2024-12
pan2024swegym
Training Software Engineering Agents and Verifiers with SWE-Gym
-
2024-05
AndroidWorld
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
-
2024-04
OSWorld
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
-
2024-03
CRADLE
CRADLE: General Computer Agents with Tool Creation and Knowledge Discovery
-
2024-03
DeepSeek-VL
DeepSeek-VL: Towards Real-World Vision-Language Understanding
-
2024-02
tang2024worldcoder
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
-
2024-01
AppWorld
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
-
2024-01
WorkArena
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
-
2023-10
SWE-bench
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
-
2023-07
WebArena
WebArena: A Realistic Web Environment for Building Autonomous Agents
-
2023-04
Generative Agents
Generative Agents: Interactive Simulacra of Human Behavior
-
2023-03
MM-ReAct
MM-ReAct: Prompting ChatGPT for Multimodal Reasoning and Action
-
2022-10
ReAct
ReAct: Synergizing Reasoning and Acting in Language Models
-
2020-10
ALFWorld
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
-
Agent RL and Continual Learning
Date
Name
Title
Paper
Github
2026-04
Skill-SD
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
-
2026-03
cursor2026realtimerl
Improving Composer through real-time RL
-
2026-03
OpenClaw-RL
OpenClaw-RL: Train Any Agent Simply by Talking
-
2026-02
xue2026acurl
Autonomous Continual Learning of Computer-Use Agents for Environment Adaptation
-
2026-02
liu2026empo2
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
-
2026-02
ye2026opcd
On-Policy Context Distillation for Language Models
-
2025-12
wei2025selfplayswerl
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
-
2025-11
MemSearcher
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
-
2025-10
wang2025agenticrlguide
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
-
2025-09
Kimi-Dev
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
-
2025-09
jin2025rluserconversations
The Era of Real-World Human Interaction: RL from User Conversations
-
2025-09
zhang2026agenticrlsurvey
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
-
2025-08
Agent Lightning
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
-
2025-08
ComputerRL
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
-
2025-05
Absolute Zero
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
-
2025-04
ReTool
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
-
2025-04
SWE-smith
SWE-smith: Scaling Data for Software Engineering Agents
-
2025-04
ToolRL
ToolRL: Reward is All Tool Learning Needs
-
2025-02
SWE-RL
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
-
2021-12
WebGPT
WebGPT: Browser-Assisted Question-Answering with Human Feedback
-
2021-01
kairouz2021advances
Advances and Open Problems in Federated Learning
-
2020-01
li2020federated
Federated Optimization in Heterogeneous Networks
-
2019-01
Towards Federated Learning at Scale
Towards Federated Learning at Scale: System Design
-
2017-06
lopezpaz2017gem
Gradient Episodic Memory for Continual Learning
-
2017-01
mcmahan2017communication
Communication-Efficient Learning of Deep Networks from Decentralized Data
-
Meta-Agents and Evolution Orchestration
Date
Name
Title
Paper
Github
2026-06
Agon
Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy
-
2026-05
Ace-Skill
Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution
-
2026-05
Continual Harness
Continual Harness: Online Adaptation for Self-Improving Foundation Agents
-
2026-05
Skill1
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning
-
2026-05
SkillOS
SkillOS: Learning Skill Curation for Self-Evolving Agents
-
2026-04
Agentic Harness Engineering
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
-
2026-04
Autogenesis
Autogenesis: A Self-Evolving Agent Protocol
-
2026-04
CORAL
CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
2026-04
Experience as a Compass
Experience as a Compass: Multi-agent RAG with Evolving Orchestration and Agent Prompts
-
2026-04
Meta-TTL
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
2026-04
pan2026mstar
M^: Every Task Deserves Its Own Memory Harness
-
2026-04
cheng2026mem2evolve
Mem^2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
-
2026-04
qiao2026mia
Memory Intelligence Agent
-
2026-04
PRIME
PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent
-
2026-04
RoboPhD
RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets
-
2026-04
yang2026memoryextraction
Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks
-
2026-04
The World Leaks the Future
The World Leaks the Future: Harness Evolution for Future Prediction Agents
-
2026-03
AgentFactory
AgentFactory: A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse
-
2026-03
AI-Supervisor
AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model
-
2026-03
ARISE
ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning
-
2026-03
AutoAgent
AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents
-
2026-03
zhang2026hyperagents
Hyperagents
-
2026-03
Memento-Skills
Memento-Skills: Let Agents Design Agents
-
2026-03
Meta-Harness
Meta-Harness: End-to-End Optimization of Model Harnesses
-
2026-03
Mimosa Framework
Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research
-
2026-03
Nurture-First Agent Development
Nurture-First Agent Development: Building Domain-Expert AI Agents Through Conversational Knowledge Crystallization
-
2026-03
RetroAgent
RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback
-
2026-03
SAGE
SAGE: Multi-Agent Self-Evolution for LLM Reasoning
-
2026-02
AOrchestra
AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
-
2026-02
Group-Evolving Agents
Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
-
2026-02
KernelBlaster
KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning
-
2026-02
xiong2026alma
Learning to Continually Learn via Meta-learning Agentic Memory Designs
-
2026-02
MemSkill
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
-
2026-02
MetaMem
MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization
-
2026-02
Position
Position: Agentic Evolution is the Path to Evolving LLMs
-
2026-02
ROMA
ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems
-
2026-02
SkillOrchestra
SkillOrchestra: Learning to Route Agents via Skill Transfer
-
2026-02
Tool-R0
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
-
2026-01
GraphPlanner
GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs
2026-01
ye2026mce
Meta Context Engineering via Agentic Skill Evolution
-
2026-01
MetaGen
MetaGen: Self-Evolving Roles and Topologies for Multi-Agent LLM Reasoning
-
2025-11
Agent0
Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
-
2025-10
MLE-Smith
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
-
2025-09
MetaEvo
MetaEvo: A Meta-Optimization Framework for Experience-Driven Agent Evolution
-
2025-08
MetaAgent
MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning
-
2025-05
AlphaEvolve
AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms
-
2025-05
dang2025evolvingorchestration
Multi-Agent Collaboration via Evolving Orchestration
-
2025-04
FlowReasoner
FlowReasoner: Reinforcing Query-Level Meta-Agents
-
2025-04
TTRL
TTRL: Test-Time Reinforcement Learning
-
2025-02
A-MEM
A-MEM: Agentic Memory for LLM Agents
-
2025-02
zhang2025maas
Multi-agent Architecture Search via Agentic Supernet
-
2024-10
AFlow
AFlow: Automating Agentic Workflow Generation
-
2024-10
AgentSquare
AgentSquare: Automatic LLM Agent Search in Modular Design Space
-
2024-08
hu2024adas
Automated Design of Agentic Systems
-
2023-05
Voyager
Voyager: An Open-Ended Embodied Agent with Large Language Models
-
Evaluation and Benchmarks
Date
Name
Title
Paper
Github
2026-05
ABRA
ABRA: Agent Benchmark for Radiology Applications
-
2026-05
Agent-BRACE
Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty
-
2026-05
wang2026dora
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
-
2026-05
raj2026consistency
Consistency as a Testable Property
-
2026-05
lee2026ctfusion
CTFusion
-
2026-05
DataClaw
DataClaw: A Process-Oriented Agent Benchmark for Exploratory Real-World Data Analysis
-
2026-05
gao2026evidencesupportedbounds
Evidence-Supported Score Bounds
-
2026-05
Evolving-RL
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
-
2026-05
surana2026gfcr
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning
-
2026-05
wu2026longmemevalv2
LongMemEval-V2
-
2026-05
MMTB
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks
-
2026-05
zhao2026rethinkingexperience
Rethinking Experience Utilization in Self-Evolving Language Model Agents
-
2026-04
agentbeats_registry2026
AgentBeats Dashboard / Agent Registry
-
2026-04
agentbeats_docs_aaa2026
Agentified Agent Assessment (AAA) & AgentBeats
-
2026-04
gurram2026agentpropbench
AgentProp-Bench
-
2026-04
ClawArena
ClawArena: Benchmarking AI Agents in Evolving Information Environments
-
2026-04
ClawBench
ClawBench: Can AI Agents Complete Everyday Online Tasks?
2026-04
EvoAgentBench
EvoAgentBench: A Multi-Domain Benchmark for Self-Evolving Agents
-
2026-04
chi2026frontiereng
Frontier-Eng
-
2026-04
SEARL
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
-
2026-04
SkillLearnBench
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
-
2026-03
agentmemorybench2026
Benchmarking Continual Agent Memory for Online Learning, Transfer, and Forgetting
-
2026-03
CUBE
CUBE: A Standard for Unifying Agent Benchmarks
-
2026-03
DomusMind
DomusMind: A Benchmark for Evaluating Lifelong Smart Home Agents Under Drift
-
2026-02
Agent World Model
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
-
2026-02
ResearchGym
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
-
2026-02
SE-Bench
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
-
2026-02
When AI Benchmarks Plateau
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
-
2026-01
DevOps-Gym
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
-
2026-01
SIP-Bench
SIP-Bench: An Open Protocol for Longitudinal Self-Improvement Evaluation
2025-07
mohammadi2025agentbenchmarking
Evaluation and Benchmarking of LLM Agents: A Survey
-
2025-07
SWE-MERA
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
-
2025-06
liu2025cer
Contextual Experience Replay for Self-Improvement of Language Agents
-
2025-05
LifelongAgentBench
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
-
2025-05
zhang2025swebenchlive
SWE-bench Goes Live!
-
2025-05
SWE-rebench
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
-
2025-03
yehudai2025agentevaluation
Survey on Evaluation of LLM-based Agents
-
2025-01
cai2025building
Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark
-
2024-12
TheAgentCompany
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
-
2024-10
MLE-bench
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
-
2024-09
wang2025awm
Agent Workflow Memory
-
2024-06
tau-bench
tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
-
2024-03
LiveCodeBench
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
-
2024-02
maharana2024locomo
Evaluating Very Long-Term Conversational Memory of LLM Agents
-
2023-08
AgentBench
AgentBench: Evaluating LLMs as Agents
-
2023-07
ToolLLM
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
-
2019-04
vandeven2019threescenarios
Three Scenarios for Continual Learning
-
Safety and Governance
Date
Name
Title
Paper
Github
2026-05
wu2026biv
Behavioral Integrity Verification for AI Agent Skills
-
2026-05
liang2026mobius
Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection
-
2026-05
LITMUS
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
-
2026-05
yin2026fate
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
-
2026-05
Proteus
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
-
2026-05
SARC
SARC: A Governance-by-Architecture Framework for Agentic AI Systems
-
2026-05
ShadowMerge
ShadowMerge: A Novel Poisoning Attack on Graph-Based Agent Memory via Relation-Channel Conflicts
-
2026-05
SkillScope
SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills
-
2026-05
SkillsVote
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
-
2026-05
STALE
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
-
2026-05
Under the Hood of SKILL.md
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
-
2026-04
AgentWatcher
AgentWatcher: A Rule-Based Prompt Injection Monitor
-
2026-04
ATBench
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
-
2026-04
Claw-Eval
Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents
-
2026-04
chang2026visualinjections
If You're Waiting for a Sign... That Might Not Be It! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems
-
2026-04
MemEvoBench
MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents
-
2026-04
SafeAgent
SafeAgent: A Runtime Protection Architecture for Agentic Systems
-
2026-04
SkillClaw
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
-
2026-04
SkillForge
SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support
-
2026-04
Sovereign Agentic Loops
Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems
-
2026-04
Spore
Spore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing
-
2026-03
From Storage to Steering
From Storage to Steering: Memory Control Flow Attacks on LLM Agents
-
2026-03
lam2026ssgm
Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework
-
2026-03
SkillProbe
SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
-
2026-03
SkillTester
SkillTester: Benchmarking Utility and Security of Agent Skills
-
2026-02
agentskills2026architecture
Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
-
2026-02
AgentSys
AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management
-
2026-02
ClawHavoc
ClawHavoc: 341 Malicious Clawed Skills Found by the Bot They Were Targeting
-
2026-02
SoK
SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
-
2026-01
WildClawBench
WildClawBench: An In-the-Wild Benchmark for AI Agents in the OpenClaw Environment
2025-12
MemoryGraft
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
-
2025-10
A-MemGuard
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
-
2025-09
InjecMEM
InjecMEM: Memory Injection Attack on LLM Agent Memory Systems
-
2025-07
shanghai2025frontierrisk
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
-
2025-07
SafeWork-R1
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45^ Law
-
2025-06
su2025autonomyrisk
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
-
2025-06
Context Manipulation Attacks
Context Manipulation Attacks: Web Agents Are Susceptible to Corrupted Memory
-
2025-06
DRIFT
DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
-
2025-06
ferrag2025promptprotocol
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
-
2025-06
RedDebate
RedDebate: Safer Responses through Multi-Agent Red Teaming Debates
-
2025-06
fang2025safemcp
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
-
2025-03
AutoRedTeamer
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
-
2024-12
Agent-SafetyBench
Agent-SafetyBench: Evaluating the Safety of LLM Agents
-
2024-12
yang2024ai45law
Towards AI-45^ Law: A Roadmap to Trustworthy AGI
-
2024-10
zhang2025asb
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-Based Agents
-
2024-03
IsolateGPT
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
-
2024-02
Agent Smith
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
-
2024-02
dong2024conversationsafety
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
-
2022-12
Constitutional AI
Constitutional AI: Harmlessness from AI Feedback
-
2022-03
ouyang2022instructgpt
Training Language Models to Follow Instructions with Human Feedback
-
Acknowledgment
This repository is maintained by the FrontisAI and Tsinghua University survey team. Its README structure follows the public awesome-list style of TsinghuaC3I/Awesome-RL-for-LRMs.
Star History
About
Awesome list and survey website for agents in the era of experience