Add support for Qwen3-4B-Thinking-2507 - #2428
Conversation
|
🤖 Hi @Rohan-Bierneni, I've received your request, and I'm working on it now! You can track my progress in the logs for more details. |
|
🤖 Hi @RissyRan, I've received your request, and I'm working on it now! You can track my progress in the logs for more details. |
It seems being throttle again: cc @richjames0 |
|
🤖 Hi @RissyRan, I've received your request, and I'm working on it now! You can track my progress in the logs for more details. |
|
Now, this seems strange: https://screenshot.googleplex.com/6UjKsrTvNQPHMxV :( |
parambole
left a comment
There was a problem hiding this comment.
Thanks for making these changes. Left a few comments.
parambole
left a comment
There was a problem hiding this comment.
LGTM. Thank you for adding this variant.
RissyRan
left a comment
There was a problem hiding this comment.
Some minor comments. Thanks!
02e427e to
36e7538
Compare
c63620f to
420a9da
Compare
Description
Maxtext already has support for Qwen3-4B architechture, but not for the -Thinking variant out of the box. This pr aims to leverage the already implemented dense Qwen3 architechture in Maxtext to add support for the Thinking variant of Qwen3-4B.
FIXES: b/440393388
Tests
Ran forward_pass_logit_checker.py on on a converted huggingface to maxtext checkpoint via the checkpoint util and achieved a acceptable average kl-divergencevalue of 0.000514: https://paste.googleplex.com/6706050174681088
Checklist
Before submitting this PR, please make sure (put X in square brackets):
gemini-reviewlabel.