Qualcomm AI Engine Direct - Scripts and accuracy improvement for Qwen3_0.6B/1.7B and Qwen 2.5_1.5B #13544

winskuo-quic · 2025-08-20T08:16:13Z

Summary

Adding static Qwen 2.5 - 1.5B to script.
Adding static Qwen 3 0.6B/1.5B to script
Adding back skip_advanced_requant.
Adding prompt + special token for calibration, which helps certain models to improve accuracy.

Example Scripts:

python examples/qualcomm/oss_scripts/llama/llama.py -b build-android -H haowhsu-linux -s 5f396958 -m SM8750 --prompt "How many r's in strawberries?" --temperature 0 --model_mode kv --max_seq_len 1024 --ptq 16a8w --decoder_model qwen3-0_6b --tasks wikitext --limit 1 --artifact ./qwen3-0_6b

Statistics on SM8750, seq_len=1024

qwen2 1.5B: ~34tok/sec. QNN on device PPL=9.4 (CPU FP=9.1)
qwen3 0.6B: ~56tok/sec. QNN on device PPL=16.8 (CPU FP=16.26)
qwen3 1.7B: ~14tok/sec. QNN on device PPL=14.1 (CPU FP=13.52)

Test plan

E2E in test_qnn_delegate.py

pytorch-bot · 2025-08-20T08:16:17Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/13544

📄 Preview Python docs built from this PR

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit dde5562 with merge base 9359481 ():

NEW FAILURE - The following job has failed:

pull / unittest-editable / linux / linux-job (gh)
exir/tests/test_remove_unused_parameters_pass.py::TestRemoveUnusedParametersPass::test_remove_unused_parameters_nested_e2e_to_edge

This comment was automatically generated by Dr. CI and updates every 15 minutes.

github-actions · 2025-08-20T08:16:49Z

This PR needs a `release notes:` label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

… 2.5 - 1.5B and Qwen 3 - 0.6B/1.7B

winskuo-quic · 2025-08-21T14:07:05Z

Hi @cccclai,
This PR is to improve accuracy for qwen3 0.6B/1.7B and enable qwen 2.5 1.5B.
The PPL score does not have much effect, which still closely align with FP CPU.
The main difference here is special tokens such as , are also feed as input during calibration, which makes the output string looks more similar to CPU FP's output.
Before, qwen3 basically could not hit eot and will repeat the same sentence. With this PR, we have tested a couple of prompts, and they can all hit eot condition.

Please have a look.
Thanks

facebook-github-bot · 2025-08-21T18:45:48Z

@cccclai has imported this pull request. If you are a Meta employee, you can view this in D80726276.

…3_0.6B/1.7B and Qwen 2.5_1.5B (pytorch#13544) ### Summary - Adding static Qwen 2.5 - 1.5B to script. - Adding static Qwen 3 0.6B/1.5B to script - Adding back `skip_advanced_requant`. - Adding prompt + special token for calibration, which helps certain models to improve accuracy. #### Example Scripts: `python examples/qualcomm/oss_scripts/llama/llama.py -b build-android -H haowhsu-linux -s 5f396958 -m SM8750 --prompt "How many r's in strawberries?" --temperature 0 --model_mode kv --max_seq_len 1024 --ptq 16a8w --decoder_model qwen3-0_6b --tasks wikitext --limit 1 --artifact ./qwen3-0_6b` #### Statistics on SM8750, seq_len=1024 qwen2 1.5B: ~34tok/sec. QNN on device PPL=9.4 (CPU FP=9.1) qwen3 0.6B: ~56tok/sec. QNN on device PPL=16.8 (CPU FP=16.26) qwen3 1.7B: ~14tok/sec. QNN on device PPL=14.1 (CPU FP=13.52) ### Test plan E2E in test_qnn_delegate.py

winskuo-quic requested a review from cccclai as a code owner August 20, 2025 08:16

meta-cla bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 20, 2025

winskuo-quic marked this pull request as draft August 20, 2025 08:51

winskuo-quic added 2 commits August 21, 2025 09:41

Qualcomm AI Engine Direct - Scripts and accuracy improvement for Qwen…

3798202

… 2.5 - 1.5B and Qwen 3 - 0.6B/1.7B

Resolve rebase conflict

dde5562

winskuo-quic force-pushed the dev1/winskuo/qwen2_5-1_5b branch from 4afc025 to dde5562 Compare August 21, 2025 14:02

winskuo-quic marked this pull request as ready for review August 21, 2025 14:02

cccclai approved these changes Aug 25, 2025

View reviewed changes

cccclai merged commit f154d50 into pytorch:main Aug 25, 2025
104 of 105 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Qualcomm AI Engine Direct - Scripts and accuracy improvement for Qwen3_0.6B/1.7B and Qwen 2.5_1.5B #13544

Qualcomm AI Engine Direct - Scripts and accuracy improvement for Qwen3_0.6B/1.7B and Qwen 2.5_1.5B #13544

Uh oh!

winskuo-quic commented Aug 20, 2025 •

edited

Loading

Uh oh!

pytorch-bot bot commented Aug 20, 2025 •

edited

Loading

Uh oh!

github-actions bot commented Aug 20, 2025

Uh oh!

winskuo-quic commented Aug 21, 2025

Uh oh!

facebook-github-bot commented Aug 21, 2025

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

Qualcomm AI Engine Direct - Scripts and accuracy improvement for Qwen3_0.6B/1.7B and Qwen 2.5_1.5B #13544

Qualcomm AI Engine Direct - Scripts and accuracy improvement for Qwen3_0.6B/1.7B and Qwen 2.5_1.5B #13544

Uh oh!

Conversation

winskuo-quic commented Aug 20, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Summary

Example Scripts:

Statistics on SM8750, seq_len=1024

Test plan

Uh oh!

pytorch-bot bot commented Aug 20, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/13544

❌ 1 New Failure

Uh oh!

github-actions bot commented Aug 20, 2025

This PR needs a release notes: label

Uh oh!

winskuo-quic commented Aug 21, 2025

Uh oh!

facebook-github-bot commented Aug 21, 2025

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

winskuo-quic commented Aug 20, 2025 •

edited

Loading

pytorch-bot bot commented Aug 20, 2025 •

edited

Loading

This PR needs a `release notes:` label