Repository navigation
v2.53.0: default GRPO-family off-policy masks, kl_clamp 10, accumulation_window loss_norm and mean-only advantages
Breaking Changes
-
default GRPO-family off-policy masks, kl_clamp 10, accumulation_window loss_norm and mean-only advantages
GRPO, GSPO and CISPO now default to
loss_norm: accumulation_window,kl_clamp: 10.0(also LLM PPO),off_policy_token_mask_bounds: [0.5, 5.0],off_policy_sequence_mask_threshold: 0.03,use_bias_correction_kl: true,adv_norm: mean_only, andtop_p: 1.0. Agents built or resumed without those keys pick up the new defaults. Opt out withkl_clamp: null,loss_norm: micro_batch,null/false/mean_std/0.95.kl_clampis a one-sided per-token bound on the K3 KL penalty: tokens further below the reference get no KL gradient. Off-policy token and sequence masks drop the policy and KL gradient of out-of-band tokens and drifted negative-advantage rows. Bias-corrected KL weights K3 by the un-clamped ratio to the old policy.The fused Liger GRPO-family loss now uses the same per-token weights as the PyTorch path under every
loss_norm. GSPO withuse_liger_loss=Trueruns (Liger sequence level) instead of raising. CISPO + Liger configs that trained atmicro_batchnow setloss_norm: accumulation_window.Arena
GRPOSpec/CISPOSpec/GSPOSpec/LLMPPOSpecexpose the new fields with those defaults.RolloutLLMSpec.top_pdefaults to 1.0.
Adaptive task sampling can pool rows into families:TaskAssigner(families=..., family_prior_strength=...)starts each row at its family's informative rate. Arena env specs exposetask_family_field(off by default) andtask_family_prior_strength(1.0).
ArenaTrainingSpec.reuse_prefix_cache_across_syncs(off by default) keeps the served LoRA name, and so vLLM's prefix cache, formax_rollout_version_lagweight versions; it requires an LLMreplay_bufferwithmax_rollout_version_lag >= 1.
Full Changelog: v2.52.0...v2.53.0