Skip to content

PTO operators binaries (per-suite, gfrun-pass) 20260904

Choose a tag to compare

@lvhao7896 lvhao7896 released this 04 Sep 08:22
· 200 commits to main since this release

PTO operators binaries (per-suite, gfrun-pass) 20260904

基于 main 分支,HEAD a0ddcc3。本 release 按 suite 拆为 2 个小包,仅含 gfrun 功能模型验证通过(PASS)的 ELF 及其反汇编 .elf.diss;FAIL/TIMEOUT 的算子不入包。compiler-workloads(coremark/dhrystone/polybench/TSVC/cTuning)已编译并用 gfrun 核验,但全部 TIMEOUT,不入包,仅在下方记录。

环境版本

组件 版本
SuperNPUBench a0ddcc3 (main)
llvm-project (clang/ld) 1ae4ee39 (dev-llvm15_56) → clang 15.0.4 (linx64v5-musl-local)
Linx-TileOP-API 804eb03 (linx) — 含未提交 template_asm.hpp(cooperative matmul LB0 = per-PE M × 4,ADR-0100 group_M 合约)
gfrun (功能模型) bc7fae00 (codex/consolidate-post-main-fixes-20260903) — dirty-built,含未提交 emulator/SoftCore.cpp(cube 布局元素索引溢出修复)
目标三元组 linx64v5-unknown-linux-musl
sysroot linx-toolchain-build/output/linx_blockisa_llvm_musl/sysroot/usr/lib/(含 libc.a/crt1.o/crti.o/crtn.o,hosted musl 链接可用)

编译器统一使用主 linx-toolchain-build 检出(COMPILER_DIR=…/linx-toolchain-build/output/linx_blockisa_llvm_musl/bin),未使用 -latest worktree 或 sibling 预编译副本。

编译命令:

# microbenchmark + one-level-arch
COMPILER_DIR=…/linx-toolchain-build/output/linx_blockisa_llvm_musl/bin
( cd microbenchmark && bash compile_all.sh all )
bash benchmark/one-level-arch/compile_all.sh
# compiler-workloads (compile-only)
python3 benchmark/compiler-workloads/run_portfolio.py --compile-only --compiler-dir "$COMPILER_DIR" --out output/compiler-workloads

gfrun 核验(micro + one-level,496 ELF,JOBS=6,单 ELF 90s 看门狗):

suite ELF 数 PASS FAIL TIMEOUT 通过率
microbenchmark 400 396 4 0 99.0%
one-level-arch 96 81 13 2 84.4%
合计 496 477 17 2 96.2%

PASS 判据:gfrun 返回码 0 + 日志出现 Reach the End of Benchmark + R2 = 0。multi_thread 算子加 -s softcore.multiThreadNum=4。

注:run_all.sh 的 tail -400 日志截断对 gfrun 输出超 400 行的 ELF(cube tmatmul fp16/i8、concat half gather)误判为 FAIL。手工复跑 5 例确认实际 PASS,已修正 summary.tsv;concat scatter 和 topk 在 300s 手工复跑仍未完成,归类为 TIMEOUT。

关键变更(vs ops-20260828)

ops-20260828(d1604cf)→ 本 release(a0ddcc3),共 10 个提交(不含中间 merge):

  • res_check 数值验证框架(72b2009):microbench bench_utils.hpp 加 golden-check,compile_all.sh 加 res_check 维度;所有 microbench 源文件加 BENCH_MAIN 宏适配。
  • fixp 新增 23 个 shared-PE 模式(9879779):fixp tmatmul 从 63→122 ELF,修复 mxacc_cscale MATH_SRC;gfrun 侧 fixpipe GroupMax/RowMaxIn 修复使 fixp FAIL 从 43→4。
  • FA persistent CUBE 重构 + nocvt→gmma 合并(99234b8/8c5c5ef):删除 fa nocvt 实现,统一为 gmma;新增 fa_fixpipe 内核(编译通过但 gfrun cooperative TMATMUL 断言待修复);matmul 新增 Sq=1024 大 shape 配置。
  • reduction CUBE_M32 行归约 subview(a244be8/4be35c6):rowsum subview 使用 logical-width 行归约 tile;gfrun 侧 CUBE subview 行归约修复使 reduction/fa/flashMLA 全恢复。
  • multi_thread 新增 8 个算子(2e53ab2):broadcast/concat/conv2d/element_wise/gather/normalization/reduction/transpose,全部 gfrun PASS。
  • deepseek CUBE 编译修复:deepseek 从 16→21 ELF,PASS 从 10→17。
  • README 刷新至 2026-09-04 回归报告(d25c7d5/a0ddcc3):验证基线 496 ELF / 96.4% 通过率,与 08-27 单一变量 +129 PASS / 0 回归。
  • gfrun 模型升级 d8903938→bc7fae00(253 commits):集中修复 CUBE subview 行归约、fixpipe GroupMax/RowMax、HiF4X2 cooperative MX、packed FP4 TCVT、cube store layout 等模型侧断言。

文件一:supernpubench_microbenchmark_20260904.tar.gz

解包后顶层为 microbenchmark/<cat>/elf/<cat>/<name>.elf[.diss]。仅含 gfrun-PASS 算子。

类别 ELF 数 PASS FAIL TIMEOUT 入包数(=PASS)
cube 11 11 0 0 11
fixp 122 118 4 0 118
memory 14 14 0 0 14
scalar 124 124 0 0 124
vector 129 129 0 0 129
合计 400 396 4 0 396

入包内容:396 个 .elf + 396 个 .elf.diss。FAIL 全部集中在 fixp 类(4 个 shared-PE tmatmul cooperative 断言),不影响其它类别。

文件二:supernpubench_one-level_20260904.tar.gz

解包后顶层为 one-level-arch/kernel/<cat>/elf/<name>.elf[.diss](multi_thread 在 one-level-arch/kernel/multi_thread/<sub>/elf/)。仅含 gfrun-PASS 算子。

类别 ELF 数 PASS FAIL TIMEOUT 入包数(=PASS)
broadcast 6 5 1 0 5
concat 4 3 0 1 3
control 6 0 6 0 0
deepseek 21 17 4 0 17
element_wise 1 1 0 0 1
fa 10 10 0 0 10
flashMLA 2 2 0 0 2
gather 1 1 0 0 1
matmul 3 3 0 0 3
multi_thread/broadcast 1 1 0 0 1
multi_thread/concat 2 2 0 0 2
multi_thread/conv2d 1 1 0 0 1
multi_thread/element_wise 1 1 0 0 1
multi_thread/fa 7 5 2 0 5
multi_thread/gather 1 1 0 0 1
multi_thread/matmul 11 11 0 0 11
multi_thread/normalization 2 2 0 0 2
multi_thread/reduction 4 4 0 0 4
multi_thread/transpose 1 1 0 0 1
multi_thread/vec 1 1 0 0 1
reduction 5 5 0 0 5
sort 1 0 0 1 0
transpose 4 4 0 0 4
合计 96 81 13 2 81

入包内容:81 个 .elf + 81 个 .elf.diss。FAIL 集中在 control(6 个 hashtable TEPL 断言)、deepseek(4 个 binary TEPL 断言)、multi_thread/fa(2 个 HIF8 convert + MXFP4 TCVT 模型侧限制);TIMEOUT 为 concat scatter 和 topk(gfrun 指令跟踪超 300s 未完成)。

compiler-workloads(不入包)

编译产出 8 个链接 ELF(coremark/dhrystone/polybench gemm+jacobi-2d/TSVC auto+off,含 2 个 .compat-linx-isa 副本)+ 206 个 .o codelet。gfrun 核验 coremark 代表性 ELF:90s 内仅 17 行输出,rc=143(TIMEOUT)。全部 cw 工作负载在 ~180K inst/s 功能模型上均无法在有限时间内完成。入包数 = 0。

说明

  • gfrun 为 dirty-built:含未提交 emulator/SoftCore.cpp(cube 布局元素索引 uint64_t 溢出修复);TileOP 含未提交 template_asm.hpp(cooperative matmul LB0 = per-PE M × 4)。无此修改 hosted musl multi_thread ELF 无法通过。
  • tarball 仅含 gfrun-PASS 的 .elf + .elf.diss;FAIL/TIMEOUT 不入包但 P/F/T 表全部公开。
  • cw freestanding 程序需 qemu 或真机运行,gfrun 功能模型不适用。
  • 排除文件类型:.o/.d/.lds/.cache/.compat-linx-isa/logs。
  • fa_fixpipe 为新内核,编译通过但 gfrun cooperative TMATMUL 断言待修复,未编入 compile.all 的 PASS 集。