PTO operators binaries (per-suite, gfrun-pass) 20260904
PTO operators binaries (per-suite, gfrun-pass) 20260904
基于 main 分支,HEAD a0ddcc3。本 release 按 suite 拆为 2 个小包,仅含 gfrun 功能模型验证通过(PASS)的 ELF 及其反汇编 .elf.diss;FAIL/TIMEOUT 的算子不入包。compiler-workloads(coremark/dhrystone/polybench/TSVC/cTuning)已编译并用 gfrun 核验,但全部 TIMEOUT,不入包,仅在下方记录。
环境版本
| 组件 | 版本 |
|---|---|
| SuperNPUBench | a0ddcc3 (main) |
| llvm-project (clang/ld) | 1ae4ee39 (dev-llvm15_56) → clang 15.0.4 (linx64v5-musl-local) |
| Linx-TileOP-API | 804eb03 (linx) — 含未提交 template_asm.hpp(cooperative matmul LB0 = per-PE M × 4,ADR-0100 group_M 合约) |
| gfrun (功能模型) | bc7fae00 (codex/consolidate-post-main-fixes-20260903) — dirty-built,含未提交 emulator/SoftCore.cpp(cube 布局元素索引溢出修复) |
| 目标三元组 | linx64v5-unknown-linux-musl |
| sysroot | linx-toolchain-build/output/linx_blockisa_llvm_musl/sysroot/usr/lib/(含 libc.a/crt1.o/crti.o/crtn.o,hosted musl 链接可用) |
编译器统一使用主 linx-toolchain-build 检出(COMPILER_DIR=…/linx-toolchain-build/output/linx_blockisa_llvm_musl/bin),未使用 -latest worktree 或 sibling 预编译副本。
编译命令:
# microbenchmark + one-level-arch
COMPILER_DIR=…/linx-toolchain-build/output/linx_blockisa_llvm_musl/bin
( cd microbenchmark && bash compile_all.sh all )
bash benchmark/one-level-arch/compile_all.sh
# compiler-workloads (compile-only)
python3 benchmark/compiler-workloads/run_portfolio.py --compile-only --compiler-dir "$COMPILER_DIR" --out output/compiler-workloads
gfrun 核验(micro + one-level,496 ELF,JOBS=6,单 ELF 90s 看门狗):
| suite | ELF 数 | PASS | FAIL | TIMEOUT | 通过率 |
|---|---|---|---|---|---|
| microbenchmark | 400 | 396 | 4 | 0 | 99.0% |
| one-level-arch | 96 | 81 | 13 | 2 | 84.4% |
| 合计 | 496 | 477 | 17 | 2 | 96.2% |
PASS 判据:gfrun 返回码 0 + 日志出现 Reach the End of Benchmark + R2 = 0。multi_thread 算子加 -s softcore.multiThreadNum=4。
注:run_all.sh 的
tail -400日志截断对 gfrun 输出超 400 行的 ELF(cube tmatmul fp16/i8、concat half gather)误判为 FAIL。手工复跑 5 例确认实际 PASS,已修正 summary.tsv;concat scatter 和 topk 在 300s 手工复跑仍未完成,归类为 TIMEOUT。
关键变更(vs ops-20260828)
ops-20260828(d1604cf)→ 本 release(a0ddcc3),共 10 个提交(不含中间 merge):
- res_check 数值验证框架(
72b2009):microbenchbench_utils.hpp加 golden-check,compile_all.sh加res_check维度;所有 microbench 源文件加BENCH_MAIN宏适配。 - fixp 新增 23 个 shared-PE 模式(
9879779):fixp tmatmul 从 63→122 ELF,修复mxacc_cscaleMATH_SRC;gfrun 侧 fixpipe GroupMax/RowMaxIn 修复使 fixp FAIL 从 43→4。 - FA persistent CUBE 重构 + nocvt→gmma 合并(
99234b8/8c5c5ef):删除 fa nocvt 实现,统一为 gmma;新增fa_fixpipe内核(编译通过但 gfrun cooperative TMATMUL 断言待修复);matmul 新增 Sq=1024 大 shape 配置。 - reduction CUBE_M32 行归约 subview(
a244be8/4be35c6):rowsum subview 使用 logical-width 行归约 tile;gfrun 侧 CUBE subview 行归约修复使 reduction/fa/flashMLA 全恢复。 - multi_thread 新增 8 个算子(
2e53ab2):broadcast/concat/conv2d/element_wise/gather/normalization/reduction/transpose,全部 gfrun PASS。 - deepseek CUBE 编译修复:deepseek 从 16→21 ELF,PASS 从 10→17。
- README 刷新至 2026-09-04 回归报告(
d25c7d5/a0ddcc3):验证基线 496 ELF / 96.4% 通过率,与 08-27 单一变量 +129 PASS / 0 回归。 - gfrun 模型升级
d8903938→bc7fae00(253 commits):集中修复 CUBE subview 行归约、fixpipe GroupMax/RowMax、HiF4X2 cooperative MX、packed FP4 TCVT、cube store layout 等模型侧断言。
文件一:supernpubench_microbenchmark_20260904.tar.gz
解包后顶层为 microbenchmark/<cat>/elf/<cat>/<name>.elf[.diss]。仅含 gfrun-PASS 算子。
| 类别 | ELF 数 | PASS | FAIL | TIMEOUT | 入包数(=PASS) |
|---|---|---|---|---|---|
| cube | 11 | 11 | 0 | 0 | 11 |
| fixp | 122 | 118 | 4 | 0 | 118 |
| memory | 14 | 14 | 0 | 0 | 14 |
| scalar | 124 | 124 | 0 | 0 | 124 |
| vector | 129 | 129 | 0 | 0 | 129 |
| 合计 | 400 | 396 | 4 | 0 | 396 |
入包内容:396 个 .elf + 396 个 .elf.diss。FAIL 全部集中在 fixp 类(4 个 shared-PE tmatmul cooperative 断言),不影响其它类别。
文件二:supernpubench_one-level_20260904.tar.gz
解包后顶层为 one-level-arch/kernel/<cat>/elf/<name>.elf[.diss](multi_thread 在 one-level-arch/kernel/multi_thread/<sub>/elf/)。仅含 gfrun-PASS 算子。
| 类别 | ELF 数 | PASS | FAIL | TIMEOUT | 入包数(=PASS) |
|---|---|---|---|---|---|
| broadcast | 6 | 5 | 1 | 0 | 5 |
| concat | 4 | 3 | 0 | 1 | 3 |
| control | 6 | 0 | 6 | 0 | 0 |
| deepseek | 21 | 17 | 4 | 0 | 17 |
| element_wise | 1 | 1 | 0 | 0 | 1 |
| fa | 10 | 10 | 0 | 0 | 10 |
| flashMLA | 2 | 2 | 0 | 0 | 2 |
| gather | 1 | 1 | 0 | 0 | 1 |
| matmul | 3 | 3 | 0 | 0 | 3 |
| multi_thread/broadcast | 1 | 1 | 0 | 0 | 1 |
| multi_thread/concat | 2 | 2 | 0 | 0 | 2 |
| multi_thread/conv2d | 1 | 1 | 0 | 0 | 1 |
| multi_thread/element_wise | 1 | 1 | 0 | 0 | 1 |
| multi_thread/fa | 7 | 5 | 2 | 0 | 5 |
| multi_thread/gather | 1 | 1 | 0 | 0 | 1 |
| multi_thread/matmul | 11 | 11 | 0 | 0 | 11 |
| multi_thread/normalization | 2 | 2 | 0 | 0 | 2 |
| multi_thread/reduction | 4 | 4 | 0 | 0 | 4 |
| multi_thread/transpose | 1 | 1 | 0 | 0 | 1 |
| multi_thread/vec | 1 | 1 | 0 | 0 | 1 |
| reduction | 5 | 5 | 0 | 0 | 5 |
| sort | 1 | 0 | 0 | 1 | 0 |
| transpose | 4 | 4 | 0 | 0 | 4 |
| 合计 | 96 | 81 | 13 | 2 | 81 |
入包内容:81 个 .elf + 81 个 .elf.diss。FAIL 集中在 control(6 个 hashtable TEPL 断言)、deepseek(4 个 binary TEPL 断言)、multi_thread/fa(2 个 HIF8 convert + MXFP4 TCVT 模型侧限制);TIMEOUT 为 concat scatter 和 topk(gfrun 指令跟踪超 300s 未完成)。
compiler-workloads(不入包)
编译产出 8 个链接 ELF(coremark/dhrystone/polybench gemm+jacobi-2d/TSVC auto+off,含 2 个 .compat-linx-isa 副本)+ 206 个 .o codelet。gfrun 核验 coremark 代表性 ELF:90s 内仅 17 行输出,rc=143(TIMEOUT)。全部 cw 工作负载在 ~180K inst/s 功能模型上均无法在有限时间内完成。入包数 = 0。
说明
- gfrun 为 dirty-built:含未提交
emulator/SoftCore.cpp(cube 布局元素索引uint64_t溢出修复);TileOP 含未提交template_asm.hpp(cooperative matmul LB0 = per-PE M × 4)。无此修改 hosted musl multi_thread ELF 无法通过。 - tarball 仅含 gfrun-PASS 的
.elf+.elf.diss;FAIL/TIMEOUT 不入包但 P/F/T 表全部公开。 - cw freestanding 程序需 qemu 或真机运行,gfrun 功能模型不适用。
- 排除文件类型:
.o/.d/.lds/.cache/.compat-linx-isa/logs。 fa_fixpipe为新内核,编译通过但 gfrun cooperative TMATMUL 断言待修复,未编入 compile.all 的 PASS 集。