Skip to content

Release v0.1.10

Choose a tag to compare

@hijkzzz hijkzzz released this 28 Sep 03:12
· 31 commits to main since this release
9644c3f

What's Changed

  • fix(geo3k): avoid duplicate turn-end tokens by @k21993 in #131
  • fix(metrics): report the IS filter fraction per sequence, not per token by @k21993 in #133
  • fix(critic): scale the MoE aux loss by the critic's own window by @k21993 in #134
  • fix(recipe): resubmit the renamed GLM-5.2 Slurm script by @MaxFreedomPollard in #135
  • fix(sft): keep VLM samples within max_len by @k21993 in #137
  • docs: polish README wording and fix stale/incorrect details by @hijkzzz in #138
  • fix(eval): count rollouts in eval_num_samples by @k21993 in #140
  • fix(geo3k): stop on nested boxed answers by @k21993 in #139
  • fix(packaging): publish an architecture-independent wheel by @k21993 in #143
  • docs(vllm): correct stale KV-cache claims around weight sync by @sysuTobo in #141
  • fix(geo3k): mark turn-limit exits as truncated by @k21993 in #142
  • docs(recipes): align MoE dispatcher comments with defaults by @k21993 in #147
  • feat(lora): LoRA for SFT and RL via AutoModel-native PEFT by @sysuTobo in #144
  • fix(agents): preserve chat rollout truncation by @k21993 in #145
  • refactor(lora): build PeftConfig once in BaseModel; trim comments; te… by @hijkzzz in #148
  • fix(agents): skip the tool on a truncated final turn; document the Ch… by @hijkzzz in #149
  • Chore/bump vllm 0.30 automodel by @hijkzzz in #150
  • Fix/review followups by @hijkzzz in #151
  • ci: on-demand two-GPU asynchronous RL end-to-end workflow by @hijkzzz in #152
  • ci(e2e-rl): run on a PR via the e2e-rl-test label; post every metric'… by @hijkzzz in #153

New Contributors

Full Changelog: v0.1.9...v0.1.10