Repository navigation
Replies: 9 comments 2 replies
WIP ProgressThe following documents the progress over the last 24 hours SetupMachine: GMKTec Evo-X2 (128GB Strix Halo) Prefill Results - t/s (higher is better)A - Current Delivery Repo
More to come... |
|
8K prefill is now around 1115t/s since my last post (Strix Halo). Given that this is on a completely mixed weight model, I think that's pretty impressive. There's other prefill areas to work on, but this MMB portion of the work is drawing to a close now. I'm now in the process of packaging the WIP work into manageable chunks, and then cross-porting what works onto GFX1201 and GFX1100. |
|
The gfx1201 porting work is a lot more involved than I had originally thought, but it is progressing well. Prefill peedups over pre-mmb baseline is around 30-40% across the board for qwen4exp (Qwen3.8-Flash-Next) and as well with some prefill speed ups for smaller quants on other MoE and dense models. There's a lot of re-deriving and optimization of quant-specific kernels. This should help out all the single-card users as well multi-card users for larger models. Closing in on a sustained 3000t/s for prefill for Qwen3.8-Flash-Next on the R9700's. I've also isolated what I believe to be a path to ~1400t/s prefill for Qwen 3.8 Flash Next for the Strix Halo for most any quant model that fits in memory, so there will be less of a need for people to restrict themselves to specially optimized quants. That work may land Wednesday. Still got gfx1100 to get through as well... |
|
Now touching on 1400t/s for prefill on the Strix Halo with Qwen3.8-Flash-Next, and ~50-65t/s for code generation. ~40-50t/s for prose. Still got a little ways to go to finish this work up, but it's now in a much better shape |
|
I was just about done, and in running a series of deep context tests, I've uncovered non-determinism at depth on the Strix Halo in the new MMB code that was adapted. Unfortunately since testing at 128K context depths isn't fast, even with ~1100-1300t/s prefill, it's likely going to take another day to nail this down. I've managed to find the area of code it's in so hopefully it should be fixed by the end of the day. Just setting expectations, this isn't going to outpace Halogen for very large prefills. Halogen runs at around 1750t/s for prefill on my box if you give it a big enough input (ie. prefill this 80K chunk of token input). For prefills <8K, my changes are faster than Halogen. The smaller the prefill, the greater the advantage my code has over Halogen. For decode speeds without MTP, Halogen is around 10-15% faster initially but the gap closes as depth grows. Halogen really are making the most of being a single-model only solution. It's hard to precisely measure what they're doing though. My solution focuses heavily on determinism and high precision accumulations. I cannot say what Halogen is doing though as we cannot see the source code, but my guess is that they're likely using smaller/faster GEMM accumulations somewhere to get that extra speed, but that's purely a guess. With MTP turned on it seems like my code is about the same sort of decode speeds as Halogen, depending on prompt specifics. Sometimes faster, sometimes slower, but overall about the same. What my code will give you though is something that's open source, and so can be debugged and iterated on by more people, and that's the ultimate goal here. Complete openness so everyone can benefit. Edit: Nailed it down to a specific feature that wasn't being correctly implemented on the MTP drafter, but is present in the main decode loop. The fix is relatively straight-forwards now for the non-determinism. It was also found that this non-determinism is NOT in the 16 patches + beta patches, but rather only in the current WIP work. |
|
Okay, this work has proven to be absolute massive, and has taken longer than I expected, but I've finished porting the work over to gfx1100 (well, as far as I can go with a single GPU), the gfx1201 work is in its closing stages and there are some VERY impressive speedups coming in for all models there. The work on gfx1100/gfx1201 porting though has exposed issues that need to be re-investigated on gfx1151 (Strix Halo), so there's another round of wrap-up needed on the Strix Halo to close those gaps, and then finally I should be able to get to the whole packaging of what is going to be an absolutely mammoth beta patch. What I'll likely do it deliver all this speed up work as an optional singular new beta patch block for people to apply as an option and test out. If, after a few days, no major issues are found, I'll likely fold that work into the main delivery patches. This is going to be a big job in itself. Thank you to those who are waiting patiently. I do believe that it'll be worth it though. In the mean time if people want to try it all out sooner, then the steps to do so are:
Doing all the above will allow those who are keen to follow the bleeding edge of my current work. I'm now seeing up to 3400t/s prefill for qwen4exp on the R9700's. Up to 1400t/s prefill for qwen4exp on Strix Halo. Prefill speeds have been boosted for most other models by 10-30% for all 3 architectures, but as always, your mileage may vary. If you see errors/issues, drop a comment here, not in the main Issues section, as this is all considered alpha-development level stuff. |
|
The WIP work has just been promoted into beta. README.md has the details, but basically check out a fresh llama.cpp Git reset to the baseline commit for the patch set, and then run llama-cpp-rdna-boosts/scripts/apply-beta.sh and that should set you up Enjoy! Edit: If you're on Strix Halo, and want the full-speed experience, ensure these environment variables are set. They trade off a small amount of accumulator accuracy for higher prefill speeds, otherwise llama-server, llama-bench, llama-cli, etc will all default to high-precision mode. Very basically, settings these forces most prefill math into BF16 mode instead of accumulating into FP32. FP32 is, of course, a little slower. These variable don't really do anything much on discrete GPUs speed-wise, but people are welcome to try them out. |
|
New baseline is cut The 28 beta patch set was then rebased onto the new release's 16-patch set and retested. Now it's time to go fix up those two issues that people found... |
|
In order to minimise my overheads, the full beta MMB campaign with all the prefill wins has now been merged into the baseline 16-patch set. There may be a number of rough edges still to sand off with wider use. This community has continued to provide excellent feedback though and I'm truly appreciative! |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
People following the repo will likely have noticed a stream of WIP updates.
These are, at least initially, focused on adapting pwilkin's excellent work for boosting Stirx Halo's prefill performance, and then generalizing it to all quantization types, as well as taking the Qwen Sparse Attention work of Qwen3.8-Flash-Next to the next level and analysing and optimizing every aspect I can.
This work is two-fold in nature. First is the GEMM optimizations for various quantization types. These boost prefill performance by approximately 15% for each of the weight quant types on RDNA3 and RDNA3.5. Early testing has shown that this does NOT seem to map well to RDNA4 as RDNA4 already has the hardware support to make the pre-existing operations faster than the RDNA3 work. I'll be re-testing this early hypothesis of course.
This quant work is universal to all model types, so it should be a flat 15% prefill speed boost on Strix Halo, with 7900XTX gains yet to be tested.
The second half of the work is speeding up Qwen Sparse Attention, which again takes inspiration from pwilkin's work as the starting point, and optimizes it even further. Already I've gotten the Strix Halo to go from a prefill of ~700t/s to ~1060t/s for mixed weight quantized models, and have the performance fall off be very shallow as context depth grows.
The QSA work should also map well to RDNA3 and RDNA4, but as hinted at above, QSA was not a dominant slowdown factor on RDNA4, so it remains to be seen how that works out there.
The work is going slowly but steadily. I hope to have something I can deliver to everyone by Tuesday.
There's still more Flash Next work planned after all of the above for which I think another 10-15% should be achievable, but again, we'll see.
One of the core aspects that's slowing this work down is ensuring that the purity guarantee is maintained. For people wondering why that is so important, it's because MTP performance hinges on the drafter and verification steps agreeing as much as possible. Now, of course, the drafter is never perfect, but if the numerics between the drafter and the batch of verifier are identical, then it's only the drafter inaccuracy that stands in the ways of better predictions, as opposed to disagreeing because the drafter and the verifier used mathematical methods to arrive at their results which potentially causes diverging token selections. This is where this repo's approach pays off, with better acceptance rates, and therefore better MTP speeds.
All reactions