| Reinforced Self-Training (ReST) for Language Modeling |
Iterative self-improvement via generating and filtering high-quality trajectories. |
Training Schemes |
Rollout & Data Strategy |
2023 |
arXiv |
| Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models |
Self-play fine-tuning against previous iterations. |
Training Schemes |
Rollout & Data Strategy |
2024 |
ICML |
| Scaling Relationship on Learning Mathematical Reasoning with Large Language Models |
Simple rejection sampling strategy for data collection. |
Training Schemes |
Rollout & Data Strategy |
2023 |
arXiv |
| DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models |
Introduces Group Relative Policy Optimization (GRPO) for group-based sampling. |
Training Schemes |
Rollout & Data Strategy |
2024 |
arXiv |
| Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning |
Down-sampling strategy to filter rollouts and reduce compute. |
Training Schemes |
Rollout & Data Strategy |
2025 |
arXiv |
| Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards |
Tree-based exploration with branching at high-uncertainty steps. |
Training Schemes |
Rollout & Data Strategy |
2025 |
arXiv |
| Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents |
Learn from failure trajectories via contrastive preference pairs. |
Training Schemes |
Rollout & Data Strategy |
2024 |
ACL |
| SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling |
Adapts soft actor-critic and hindsight replay for open-ended exploration. |
Training Schemes |
Rollout & Data Strategy |
2024 |
arXiv |
| Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning |
Reverse curriculum learning starting from goal proximity. |
Training Schemes |
Rollout & Data Strategy |
2024 |
ICML |
| Voyager: An Open-Ended Embodied Agent with Large Language Models |
Uses predefined progressions in open-ended worlds (Minecraft). |
Training Schemes |
Rollout & Data Strategy |
2023 |
arXiv |
| RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments |
Dynamically generates tasks based on current agent performance. |
Training Schemes |
Rollout & Data Strategy |
2025 |
arXiv |
| WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning |
Self-correcting curriculum for web agents. |
Training Schemes |
Rollout & Data Strategy |
2025 |
ICLR |
| VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models |
Uses variance of group rewards to prioritize medium-difficulty tasks. |
Training Schemes |
Rollout & Data Strategy |
2025 |
arXiv |
| Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning |
Uses exact matching outcome rewards for definitive tasks. |
Training Schemes |
Feedback & Credit |
2025 |
COLM |
| ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning |
Exact matching outcome rewards for search tasks. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis |
Functional verification for open-ended intent satisfaction. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning |
Optimizes efficiency by penalizing retrieval costs. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning |
Reward modeling for open-ended search tasks. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning |
Intent satisfaction rewards for search agents. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent |
Optimizes efficiency and functional verification. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| DeepRAG: Thinking to Retrieve Step by Step for Large Language Models |
Optimizes efficiency by explicitly penalizing unnecessary retrieval actions. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| UR^2: Unify RAG and Reasoning through Reinforcement Learning |
Efficiency-aware reward modeling for retrieval. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| ReZero: Enhancing LLM search ability by trying one-more-time |
Rewards focused on relevance and formatting in IR. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| s3: You Don't Need That Much Data to Train a Search Agent via RL |
IR reward modeling for query diversity and relevance. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning |
Planning-centric rewards for search agents. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| DeepDiver: Adaptive Search Intensity Scaling via Open-Web Reinforcement Learning |
IR rewards for deep exploration of topics. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| O^2-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering |
Outcome-oriented rewards for search quality. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning |
Tool-augmented reward modeling for long-form agentic tasks. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning |
Uses Shortest Path Reward Estimation for trajectory quality. |
Training Schemes |
Feedback & Credit |
2025 |
NeurIPS |
| Synthetic Data Generation and Multi-Step Reinforcement Learning for Reasoning and Tool Use |
Holistic history analysis for process rewards. |
Training Schemes |
Feedback & Credit |
2025 |
COLM |
| Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers |
Process rewards for search exploration steps. |
Training Schemes |
Feedback & Credit |
2025 |
NeurIPS |
| ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs |
Dense supervision via process reward models (PRM). |
Training Schemes |
Feedback & Credit |
2025 |
NeurIPS |
| Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning |
Granular evaluation of specific reasoning steps. |
Training Schemes |
Feedback & Credit |
2025 |
NeurIPS |
| Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation |
Verification-based process rewards for reasoning. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents |
Process supervision for verifiable meta-reasoning steps. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization |
Intermediate search quality evaluation. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Coordinating Search-Informed Reasoning and Reasoning-Guided Search in Claim Verification |
Hierarchical rewards for information seeking. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning |
Rewards parallel decomposition efficiency. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment |
Uses LLMs to grade/critique intermediate steps. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Advancing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization |
Collaborative grading for UI actions. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic |
LLM-based critique for search trajectories. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models |
Verbal feedback for self-correction via retrospective analysis. |
Training Schemes |
Feedback & Credit |
2025 |
NeurIPS |
| Reflexion: language agents with verbal reinforcement learning |
Verbal reinforcement for iterative refinement. |
Training Schemes |
Feedback & Credit |
2023 |
NeurIPS |
| VinePPO: Refining Credit Assignment in RL Training of LLMs |
Statistical credit assignment using rollout branching. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Exploiting Tree Structure for Credit Assignment in RL Training of LLMs |
Tree-structured estimation for step-wise credit. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution |
Trains dedicated value networks for step influence. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Agentic Reinforcement Learning with Implicit Step Rewards |
Implicit value prediction for self-taught reasoners. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| GRPO-$\lambda$: Credit Assignment improves LLM Reasoning |
Reformulated objective for granular updates without value heads. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Group-in-Group Policy Optimization for LLM Agent Training |
Group-in-Group policy optimization for credit assignment. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design |
Multi-turn adaptation of GRPO credit assignment. |
Training Schemes |
Feedback & Credit |
2025 |
arXiv |
| Proximal Policy Optimization Algorithms |
Standard clipped gradient optimization (Trust Region). |
Training Schemes |
Policy Optimization |
2017 |
arXiv |
| VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks |
Variant of PPO adapted for agentic stability. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning |
Scaled GRPO for reasoning tasks. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| DAPO: An Open-Source LLM Reinforcement Learning System at Scale |
Systematized GRPO for scale and distributed training. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| URPO: A Unified Reward & Policy Optimization Framework for Large Language Models |
Unifies policy optimization with reward modeling. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling |
Geometry-aware objectives for tree-structured policies. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning |
Lightweight PPO with entropy bonuses. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems |
Survey on reasoning degeneracy and regularization. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning |
PRM-free step-level estimation for efficiency. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards |
Asymmetric REINFORCE for off-policy data reuse. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts |
Selective rollout strategy to optimize compute budget. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| A Survey on the Optimization of Large Language Model-based Agents |
Survey on efficient optimization strategies. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning |
Uses early stopping triggers (KL spikes) for stability. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios |
Overview of reasoning stability and hacking. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement |
Multi-objective scaffolds for safety and performance. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| Safe RLHF: Safe Reinforcement Learning from Human Feedback |
Constrained MDP formulation for safety. |
Training Schemes |
Policy Optimization |
2023 |
ICLR 2024 |
| SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning |
Safety constraints for vision-language agents. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| Agentic Reinforcement Learning for Search is Unsafe |
Penalizes harmful queries in search agents. |
Training Schemes |
Policy Optimization |
2025 |
arXiv |
| MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory |
Fine-tunes models to execute explicit read/write API calls. |
Training Schemes |
Training-Time Memory |
2024 |
arXiv |
| MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation |
Trains multi-stage summarization pipelines. |
Training Schemes |
Training-Time Memory |
2023 |
arXiv |
| Augmenting Language Models with Long-Term Memory |
Introduces decoupled memory encoders. |
Training Schemes |
Training-Time Memory |
2023 |
NeurIPS |
| Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection |
Trains reflection tokens to trigger on-demand retrieval. |
Training Schemes |
Training-Time Memory |
2024 |
ICLR |
| MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents |
Uses RLVR to compress context into a constant footprint. |
Training Schemes |
Training-Time Memory |
2025 |
arXiv |
| MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent |
Adapts DAPO for streaming document processing. |
Training Schemes |
Training-Time Memory |
2025 |
arXiv |
| Memory-R1: Enhancing large language model agents to manage and utilize memories via reinforcement learning |
Uses PPO/GRPO to train a dedicated memory-manager agent. |
Training Schemes |
Training-Time Memory |
2025 |
arXiv |
| LongMemEval: Benchmarking chat assistants on long-term interactive memory |
Benchmark for knowledge updates and abstention. |
Training Schemes |
Training-Time Memory |
2024 |
arXiv |
| LongBench-v2: Towards deeper understanding and reasoning on realistic long-context multitasks |
Benchmark for extreme-context (2M tokens) tasks. |
Training Schemes |
Training-Time Memory |
2025 |
ACL |