0.7.0
CUDA-Q QEC 0.7.0
This is a release of CUDA-Q QEC version 0.7.0. (CUDA-Q Solvers is not included in this release.)
CUDA-Q QEC 0.7.0 is a decoder-focused release that makes the decoder stack faster, more scalable under real-time constraints, and easier to extend. The most notable changes are:
- Lower logical error rates under real-time deadlines. The closed-source
nv-qldpc-decoderadds a new gamma-ensemble Relay-BP mode (contributed by @kvmto) that runs several independent gamma trajectories ("lanes") in parallel on a single GPU and exits as soon as any lane converges. By contracting the decoder's latency tail — what matters most when every decode must finish within a fixed wall-clock budget, as in real-time QEC — it lowers the logical error rate under fixed decoding deadlines by up to ~89× for large bivariate-bicycle qLDPC codes on a GB200 (for example ~89× for[[288,12,18]]and ~41× for[[144,12,12]]at deadlines of ~1–5 ms), while also tightening worst-case (p99.99) latency by ~2.7-5.5×. See Improving Relay BP Decoding With Gamma Ensembles for more details. - Sparse parity-check matrices end to end, so large qLDPC codes no longer need their parity-check matrices materialized as dense tensors — including native
scipy.sparseinput for thenv-qldpc-decoder. - Decoder construction directly from Stim detector-error-model (DEM) strings, plus a new DEM-native Chromobius color-code decoder.
- A GPU/CPU
dem_samplingcapability - Surface-code orientation (XV/XH/ZV/ZH) control.
- A declarative decoder-configuration schema that lets third-party decoders be fully YAML-configurable from their own shared library with no CUDA-Q QEC rebuild.
Beyond gamma-ensemble, the nv-qldpc-decoder also gains sum-product BP variants and on-device observables output, along with several performance and correctness fixes. Under the hood, the Python bindings were migrated from Pybind11 to Nanobind for upstream CUDA-Q compatibility, and the decoder/realtime code paths were decoupled from the CUDA-Q runtime (including a dedicated CUDA-Q QEC logger).
Please check out the docs and examples for how to get started using CUDA-Q QEC.
Dependency note: CUDA-Q QEC 0.7.0 depends on CUDA-Q 0.15.1 and builds against its published images (CUDA 12.6 and 13.0).
Realtime decoding: The real-time GPU decoding capabilities introduced in 0.6.0 (the CUDA-Q Realtime
HOST_LOOPbridge and the Relay-BP / PyMatching predecoder examples) continue to work in 0.7.0.
Features and Enhancements (QEC) 🎉
Decoders & detector error models
- Sparse parity-check matrix support for decoders by @vedika-saravanan in #550
- Sparse-aware PCM utility migration by @vedika-saravanan in #602
- Make canonicalization a member of
sparse_binary_matrixby @vedika-saravanan in #599 - Adopt
scipy.sparseas optional interop by @bmhowe23 in #590 - Fix dense → sparse conversion in
get_decoderto avoid redundant copies by @bmhowe23 in #589 - Support Stim DEM strings in
get_decoderby @vedika-saravanan in #571 - DEM from Stim text can use error decompositions by @eliotheinrich in #615
- Add
dem_samplingwith CPU and GPU backends (C++ and Python) by @kvmto in #479 - Add Chromobius decoder to the decoder plugins by @wsttiger in #546
- Add YAML and Python config support for Chromobius / TRT global decoder by @melody-ren in #633
- Extend
trt_decoderwith global decoder chaining by @bmhowe23 in #524 - Expand TRT decoder YAML config for composite decoding by @wsttiger in #536
- Add CLI override flags and
from_nametoPipelineConfigby @wsttiger in #503 - Fix
ai_decoder_serviceTRT builder for quantized ONNX (FP8) by @wsttiger in #507 - Add boundary-aware overloads for
canonicalize_for_roundsand sliding-window decoder by @eliotheinrich in #656 - Use canonical soft-to-hard decoder conversion by @melody-ren in #553
- Return Python decode result as NumPy arrays by @melody-ren in #558
trt_decodernow throws on inference failure instead of returning stale/zeroed results by @melody-ren in #680
Note (composite decoding): The
trt_decodercan chain a second-stage "global" decoder (for example PyMatching or Chromobius) via theglobal_decoderandglobal_decoder_paramsoptions. When constructing the decoder directly (rather than from a YAML config, which fills this in automatically), you must supplyglobal_decoder_params— an empty map is fine — wheneverglobal_decoderis set, or the global stage is skipped.
Codes & circuits
- Add surface-code orientation support (XV/XH/ZV/ZH) by @melody-ren in #637
- Add detector support to memory circuits by @eliotheinrich in #636
Extensibility & diagnostics
- Declarative decoder parameter schemas: pluggable realtime decoder configuration by @bmhowe23 in #679
- Create CUDA-Q QEC logger by @tlshannon in #630
- Add
cuda_device_idplacement knob for GPU decoders by @melody-ren in #690 - Route every decoder device pin through one resolver and two wrappers by @melody-ren in #698
Realtime decoding infrastructure (experimental)
- Standalone realtime QEC decoding server by @bmhowe23 in #666
- Add QEC decoder-server core and CQR adapter by @vedika-saravanan in #653
- Add host-side in-process-RPC path for real-time QEC decoding by @cketcham2333 in #609
- Add PyMatching HOST_CALL decoder-server RPC path by @cketcham2333 in #600
- Add vanilla PyMatching support to realtime decoder config by @vedika-saravanan in #614
- qec/realtime: device-graph scheduler for per-round Hololink QLDPC decoding by @cketcham2333 in #631
- Add bundled
decoder_contextstruct for measurement extraction by @eliotheinrich in #671 - Virtualize methods in realtime decoder API by @bmhowe23 in #674
- Consolidate decoder RPC wire format into a single header by @bmhowe23 in #681
- Move nv-qldpc-decoder schema registration into its plugin by @bmhowe23 in #701
- Decoding server: add virtual hooks for decoder plugins to set
D/Osparse matrices by @tlshannon in #746
Realtime infrastructure — fixes & test coverage (experimental):
- Fix GB200
gpu_rocedecoding-server validation issues by @vedika-saravanan in #683 - Add QEC
device_callhost-dispatch E2E coverage by @vedika-saravanan in #628 - Relay BP (nv-qldpc)
gpu_roceprofile for the HSB decoding-server test by @cketcham2333 in #670 - Add a TRT + PyMatching decoder profile to the HSB FPGA decoding-server test by @cketcham2333 in #673
- Fix CQR surface-code test to use Relay-BP by @vedika-saravanan in #691
nv-qldpc-decoder Updates (Closed Source)
New features and options
- Gamma-ensemble sequential relay-BP (@kvmto) — a new sparse-GPU kernel adds a
gamma_ensemble_sizeoption (1/2/4/8; default 1 = disabled) that runs multiple parallel gamma "lanes" per relay iteration with race-to-fastest semantics (first lane to satisfy the stopping criterion wins; ties broken by lowest-weight correction). Supported on the sparse-GPU single-decode path withcomposition=1andbp_method=3or5. - Sum-product BP variants (@bmhowe23) —
bp_methodgains4(sum-product + memory) and5(sum-product + damped memory), both requiringuse_sparsity=True. Sequential relay (composition=1) now acceptsbp_method=5, withgamma0,gamma_dist, andexplicit_gammasextended to the new methods. - Native
scipy.sparseparity-check-matrix input (@vedika-saravanan, @bmhowe23) — the decoder was migrated to thesparse_binary_matrixAPI and accepts anyscipy.sparseformat (CSR/CSC/COO/…) directly, with no dense.toarray()/.todense()conversion. - On-device observables output (@melody-ren) — a new optional
Omatrix (shapenum_observables × block_size) makesdecode()/decode_batch()return observable flips (O · correction mod 2) directly instead of the raw correction vector.
Correctness fixes
- Fixed a race condition and undersized allocations in the batched and persistent-buffer sparse-GPU BP paths (@bmhowe23).
- Fixed a sparse batched-GPU offset overflow and an out-of-bounds access in the OSD solver (
osd_solver_gf2) (@melody-ren).
Build / packaging
- AArch64 builds now target
-march=armv8-a(was-march=native) for portability; x86-64 remains at x86-64-v3 (AVX2/SSE2). - Wired the plugin into the new declarative decoder-config schema.
Bug Fixes (QEC) 🐛
- Fix non-physical
TensorNetworkDecoderposteriors by @vedika-saravanan in #707 - Fix PyMatching decoding for observable indices above 31 by @kaiqiy-nv in #697
- Fix surface-code hook errors with an orientation-aware CNOT schedule by @bmhowe23 in #672
- Fix Python QEC conversion for empty repetition-code X matrices by @kaiqiy-nv in #660
- Stabilizer round kernels in Python no longer fail when the return-type annotation is incorrect by @eliotheinrich in #642
- Fix OOB r/w when batch failure returns empty results by @kaiqiy-nv in #625
- Fix DEM canonicalization to preserve the modeled error process by @bmhowe23 in #610
- Throw an error if no noise is provided to DEM Python binding functions by @kaiqiy-nv in #565
- Fix typo on pointer access by @kaiqiy-nv in #535
- Fix potential use-after-free and add checks before heavy calculations by @kaiqiy-nv in #500
Breaking Changes & Deprecations ⚠️
- Sparse parity-check matrix decoder API (#550): decoder constructors now take
const sparse_binary_matrix &. Dense callers keep working via implicit conversion, but out-of-tree decoder plugins must be recompiled. - Detector-based memory circuits (#636):
sample_memory_circuitsyndrome output shape changes from(nShots·nRounds) × nAncillatonShots × nDetectors(Stim-aligned). Addsx/z_sample_memory_circuitand adecompose_errorsflag ondem_from_memory_circuit. - NumPy decode results (#558):
DecoderResult.result(anddecode_batch, async, and tuple-unpacking variants) now return a 1-D NumPy array instead of a Python list. - Pybind → Nanobind (#512): Python bindings migrated to Nanobind (for CUDA-Q compatibility) and the cuquantum-python dependency was bumped. A build/packaging change for downstream consumers.
- Revert noise-learning support (#714): reverts the differentiable noise-learning (NMOptimizer) support for the TN decoder that was introduced in #526 and #577. Do not rely on that API in 0.7.0.
- Remove
hololink_gpu_idenv var (#692): removed in favor of thecuda_device_idplacement knob;reconcile_gpu_roce_devicerenamed toresolve_decode_device. - Remove
--enable-mlirbuild flag (#586) and remove usage of CUDA-Q library mode (#523). - Rename
CUDAQ_REGISTER_TYPE→CUDAQ_EXT_PT_REGISTER_TYPE(#542).
Documentation ✏️
- Document sparse parity-check matrix support by @vedika-saravanan in #657
- Add docs and example for decoder construction from Stim DEM strings by @vedika-saravanan in #598
- Update
trt_decoderdocumentation for recent enhancements by @wsttiger in #408 - Add Predecoder Decoding with CUDA-Q Realtime by @wsttiger in #497
- Add docs for the realtime FPGA-based predecoder + PyMatching demo by @wsttiger in #498
- Add Relay BP Decoding with CUDA-Q Realtime by @cketcham2333 in #488
- Update sliding-window documentation by @cketcham2333 in #421
- Add surface code to pre-built QEC code docs by @vedika-saravanan in #495
- Update nv-qldpc-decoder docs by @bmhowe23 in #499
- Document how to obtain the nv-qldpc-decoder plugin by @cketcham2333 in #532
- Remove
automoduleexclude-members forcudaq_qecPython API by @vedika-saravanan in #513 - Surface code demo 1 refactor by @eliotheinrich in #685 (follow-up in #712)
- Docs: Fix Sphinx warnings by @tlshannon in #708
- Docs: Set warnings as errors by @tlshannon in #715
- Add PR guidelines and checklist by @bmhowe23 in #555
- Document PyMatching/Chromobius decoders, surface-code
orientation,cuda_device_id, and the NumPy decode-result change by @bmhowe23 in #744 - Add
dem_samplingdocumentation by @kvmto in #511 - Update sliding-window and measurement-to-detector docs by @eliotheinrich in #728
- Restore and document v0.7 realtime
HOST_LOOPsupport by @vedika-saravanan in #743 - Update
nv-qldpc-decoderdocs for 0.7.0 (gamma-ensemble, sum-product BP,Oobservables,repeatable, sparse-PCM input) by @bmhowe23 in #753 - Add gamma-ensemble Relay-BP performance guide (Improving Relay BP Decoding With Gamma Ensembles) by @eliotheinrich in #755
Common / Misc
CUDA-Q decoupling & dependency cleanup
- Remove CUDA-Q dependencies in QEC decoders and realtime by @tlshannon in #645
- Remove CUDA-Q deps: additional server cmake closure tests by @tlshannon in #659
- Tensor Network dependency cleanup by @vedika-saravanan in #522
- Prevent vendored dependency artifacts leakage by @vedika-saravanan in #601
- Guard deprecated TRT FP16 builder API behind
NV_TENSORRT_MAJOR < 11by @bmhowe23 in #641
CUDA-Q version alignment (LLVM 22 / namespace unification / bumps)
- Update libraries for upstream CUDA-Q changes (LLVM 22) by @bmhowe23 in #531
- Fix building Python wheels with LLVM 22 by @anjbur in #540
- Update clang-format for LLVM 22 by @anjbur in #549
- Adapt to CUDA-Q internal-namespace unification (
cudaq::detail) by @khalatepradnya in #583 - Numerous CUDA-Q commit/version bumps culminating in the 0.15.1 target (#688, #675, #654, #640, #596, #588, #580, #578, #574, #551, #545, #527, #521, #516, and the release pin in #719 / "Target cuda-quantum 0.15.1")
CI / images / wheels / build plumbing
- Use Hugging Face to fetch the Ising predecoder artifact by @melody-ren in #713
- Create new workflow for multi-platform decoding server image by @anjbur in #700
- Dev image updates for CMake 4 compatibility by @anjbur in #689
- Add realtime build to cudaqx-dev image by @anjbur in #662
- Add realtime to cached CUDA-Q build by @anjbur in #624
- Add wheel build against latest CUDA-Q to nightly workflow by @anjbur in #616
- Add nightly build using latest cuda-quantum main by @anjbur in #569
- Add wheel build and test to CI PR checks by @anjbur in #581
- ci: enable NUMA affinity syscalls and add gated multi-GPU QEC test lane by @kvmto in #648
- Reduce size of cudaqx-dev image and make TensorRT dependency checking automatic by @bmhowe23 in #518
- Upgrade
prune_cudaqx-dev_by_sha.shto OCI-aware pruning by @bmhowe23 in #584 - Enable docs deployment and include the docs build in PR verification by @anjbur in #721 and #741
- (Plus additional CI/wheel cache and workflow fixes: #742, #700, #669, #635, #638, #608, #593, #597, #594, #585, #561, #570, #567, #562, #560, #544, #541, #537, #502)
Tests / coverage
- Add unit tests for QEC by @kaiqiy-nv in #605
- New DEM canonicalization tests by @kaiqiy-nv in #568
- QEC GPU Python tests (follow-up to #479) by @kvmto in #510