Replies: 2 comments
|
Thanks for asking before building, @Ru0k3; for something this size that is exactly the right order. A new, non-generative training task ( The same applies to #1187: the review there stands, and it needs a tracking issue with the scope agreed before the |
|
@Ru0k3, the decision on the scope question: not now. A native Thank you for writing up the feasibility pass before building it; it made this an easy question to answer, and it will be the starting point when this comes back. |
Uh oh!
There was an error while loading. Please reload this page.
Hi — I've been using Soup for RL fine-tuning and wanted to float something before building more of it than makes sense.
I opened #1187 adding
training.laya_reward, which lets a Laya/Jev-style typed-decision model (score/noul answers, not generated text) be used as a GRPO/PPO reward source. That's a small addition using Soup's existingreward_fnseam.The bigger question is whether Soup should go further and let people train a typed-decision model natively —
task: decision_model— instead of needing Laya's separate Kaggle notebook. I've done a source-level feasibility pass on both repos and wanted to check in before writing more code, because this is a real architectural departure, not a small extension:AutoModel), not autoregressive token logits — closer to Soup'sembeddingtask than togrpo/ppo.trl.GRPOTrainer's completion-based reward loop.state+ typedquestions+golddistributions), its own collator, and a post-training calibration-fitting step.I'd also plan to preserve Laya's native checkpoint layout exactly, so anything trained this way would drop straight into #1187's
laya_rewardadapter with no bridging code needed.Given the scope, this isn't one PR — it'd need to land as 3–4 smaller ones (checkpoint/data loading → training loop → calibration → CLI dispatch), each independently reviewable.
Before I put more time in: is this in scope for Soup at all, or do you see non-generative/decision-model training as outside what this project wants to own? And if it's in scope, would you rather see it as a first-class
task:, or is some kind of plugin/extension point a better fit long-term?Happy to share the full feasibility writeup if useful.
All reactions