Concurrent MLX GPU + Apple Neural Engine serving for Laya — looking for cross-SoC benchmarks #4551
tc3oliver
started this conversation in
Show and tell
Replies: 1 comment
|
Update: the first external cross-SoC result has landed. An M4 Pro (48 GB, macOS 27.0) passed both MLX FP16 and ANE parity with 0 hard mismatches. After local calibration, short single-question requests routed to the ANE, and the heterogeneous GPU+ANE check also passed. This is one community Result: Still looking for M1/M2/M3/M5 and other Apple Silicon results: |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
laya-appleis a correctness-validated heterogeneous Laya runtime for Apple silicon. It serves requests concurrently across the MLX GPU and the Apple Neural Engine.Same 102 requests, same arrival times — GPU-only vs concurrent MLX GPU + Apple Neural Engine serving.
1080p MP4 · where every number on screen comes from. The workload is laya-typed-decisions with open-loop bursty arrivals, and playback is 3.4× slower than real time.
What the video shows:
Key benchmark
Mixed-workload throughput, GPU-only against concurrent GPU + ANE serving:
Laya(execution="workers")instance.benchmarks/v1.0.md.The gain comes from using both engines concurrently, not from the ANE being 4.57× faster per request. These are results from one machine.
Correctness finding
In the tested configuration, the ordinary Core ML export ran on the ANE without errors. On some workloads, though, it changed decisions against upstream PyTorch: up to 85 hard mismatches on the validation rows. The same export on
CPU_AND_GPUmatched. This finding covers these models on this setup, not Core ML or the ANE in general.laya-apple therefore uses only ANE artifacts that pass two checks: a parity gate and a placement validation. On the shipped validation rows, the laya-apple ANE path has 0 hard decision mismatches (
docs/correctness.md).Why heterogeneous routing
Help test other Apple silicon
I only have an M4 Max, so I don't want to assume these routing thresholds or gains generalize across Apple silicon. If you have another Mac, a result from it is the most useful contribution right now:
No code changes required. After setup, run:
The script writes a
hardware-results/bundle that you can submit as-is in a pull request. Each issue has the full steps. Results that differ from mine, including failures, are just as useful.Repository: https://github.com/tc3oliver/laya-apple
Community benchmark matrix: https://github.com/tc3oliver/laya-apple/blob/main/docs/community-benchmarks.md
Reproducible benchmark: https://github.com/tc3oliver/laya-apple/blob/main/benchmarks/v1.0.md
Listed in Laya's upstream Community Tools: https://github.com/NandhaKishorM/laya#community-tools
All reactions