Source code for ARMOR (Anchor Rollout and Mixed Optimization for RL) https://arxiv.org/abs/2607.10481
Please follow the DAPO's recipe and QAE to prepare the data and models, then you can run the scripts to start continual training.
The continual training scripts are provided in scripts/.