Field report: 0.6.0 → 0.6.2 validation numbers from a 24/7 production fleet (M3 Ultra + M4 Max) #2862
anicaise-ai
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Some validation datapoints from our fleet after upgrading, in case they're useful:
v0.6.2 vs v0.6.1, Qwen3.6-35B-A3B-mxfp8 (MTP on, hybrid GDN), M4 Max 128 GB, A-B-A protocol (0.6.2 → 0.6.1 → 0.6.2, fresh unique long prompts per phase to defeat the persistent SSD prefix cache):
No regression anywhere; 0.6.2 is our production build on this host now.
Earlier confirmations of your announced gains, for the record:
reasoning_contentseparation eliminated a whole class of "reasoning fragments reaching users" bugs for us on complete responses (see separate issue for the finish=length edge case).All reactions