Skip to content

0.7.0

Choose a tag to compare

@bmhowe23 bmhowe23 released this 03 Aug 19:02
· 3 commits to releases/v0.7.0 since this release
02c4339

CUDA-Q QEC 0.7.0

This is a release of CUDA-Q QEC version 0.7.0. (CUDA-Q Solvers is not included in this release.)

CUDA-Q QEC 0.7.0 is a decoder-focused release that makes the decoder stack faster, more scalable under real-time constraints, and easier to extend. The most notable changes are:

  • Lower logical error rates under real-time deadlines. The closed-source nv-qldpc-decoder adds a new gamma-ensemble Relay-BP mode (contributed by @kvmto) that runs several independent gamma trajectories ("lanes") in parallel on a single GPU and exits as soon as any lane converges. By contracting the decoder's latency tail — what matters most when every decode must finish within a fixed wall-clock budget, as in real-time QEC — it lowers the logical error rate under fixed decoding deadlines by up to ~89× for large bivariate-bicycle qLDPC codes on a GB200 (for example ~89× for [[288,12,18]] and ~41× for [[144,12,12]] at deadlines of ~1–5 ms), while also tightening worst-case (p99.99) latency by ~2.7-5.5×. See Improving Relay BP Decoding With Gamma Ensembles for more details.
  • Sparse parity-check matrices end to end, so large qLDPC codes no longer need their parity-check matrices materialized as dense tensors — including native scipy.sparse input for the nv-qldpc-decoder.
  • Decoder construction directly from Stim detector-error-model (DEM) strings, plus a new DEM-native Chromobius color-code decoder.
  • A GPU/CPU dem_sampling capability
  • Surface-code orientation (XV/XH/ZV/ZH) control.
  • A declarative decoder-configuration schema that lets third-party decoders be fully YAML-configurable from their own shared library with no CUDA-Q QEC rebuild.

Beyond gamma-ensemble, the nv-qldpc-decoder also gains sum-product BP variants and on-device observables output, along with several performance and correctness fixes. Under the hood, the Python bindings were migrated from Pybind11 to Nanobind for upstream CUDA-Q compatibility, and the decoder/realtime code paths were decoupled from the CUDA-Q runtime (including a dedicated CUDA-Q QEC logger).

Please check out the docs and examples for how to get started using CUDA-Q QEC.

Dependency note: CUDA-Q QEC 0.7.0 depends on CUDA-Q 0.15.1 and builds against its published images (CUDA 12.6 and 13.0).

Realtime decoding: The real-time GPU decoding capabilities introduced in 0.6.0 (the CUDA-Q Realtime HOST_LOOP bridge and the Relay-BP / PyMatching predecoder examples) continue to work in 0.7.0.


Features and Enhancements (QEC) 🎉

Decoders & detector error models

  • Sparse parity-check matrix support for decoders by @vedika-saravanan in #550
  • Sparse-aware PCM utility migration by @vedika-saravanan in #602
  • Make canonicalization a member of sparse_binary_matrix by @vedika-saravanan in #599
  • Adopt scipy.sparse as optional interop by @bmhowe23 in #590
  • Fix dense → sparse conversion in get_decoder to avoid redundant copies by @bmhowe23 in #589
  • Support Stim DEM strings in get_decoder by @vedika-saravanan in #571
  • DEM from Stim text can use error decompositions by @eliotheinrich in #615
  • Add dem_sampling with CPU and GPU backends (C++ and Python) by @kvmto in #479
  • Add Chromobius decoder to the decoder plugins by @wsttiger in #546
  • Add YAML and Python config support for Chromobius / TRT global decoder by @melody-ren in #633
  • Extend trt_decoder with global decoder chaining by @bmhowe23 in #524
  • Expand TRT decoder YAML config for composite decoding by @wsttiger in #536
  • Add CLI override flags and from_name to PipelineConfig by @wsttiger in #503
  • Fix ai_decoder_service TRT builder for quantized ONNX (FP8) by @wsttiger in #507
  • Add boundary-aware overloads for canonicalize_for_rounds and sliding-window decoder by @eliotheinrich in #656
  • Use canonical soft-to-hard decoder conversion by @melody-ren in #553
  • Return Python decode result as NumPy arrays by @melody-ren in #558
  • trt_decoder now throws on inference failure instead of returning stale/zeroed results by @melody-ren in #680

Note (composite decoding): The trt_decoder can chain a second-stage "global" decoder (for example PyMatching or Chromobius) via the global_decoder and global_decoder_params options. When constructing the decoder directly (rather than from a YAML config, which fills this in automatically), you must supply global_decoder_params — an empty map is fine — whenever global_decoder is set, or the global stage is skipped.

Codes & circuits

Extensibility & diagnostics

  • Declarative decoder parameter schemas: pluggable realtime decoder configuration by @bmhowe23 in #679
  • Create CUDA-Q QEC logger by @tlshannon in #630
  • Add cuda_device_id placement knob for GPU decoders by @melody-ren in #690
  • Route every decoder device pin through one resolver and two wrappers by @melody-ren in #698

Realtime decoding infrastructure (experimental)

  • Standalone realtime QEC decoding server by @bmhowe23 in #666
  • Add QEC decoder-server core and CQR adapter by @vedika-saravanan in #653
  • Add host-side in-process-RPC path for real-time QEC decoding by @cketcham2333 in #609
  • Add PyMatching HOST_CALL decoder-server RPC path by @cketcham2333 in #600
  • Add vanilla PyMatching support to realtime decoder config by @vedika-saravanan in #614
  • qec/realtime: device-graph scheduler for per-round Hololink QLDPC decoding by @cketcham2333 in #631
  • Add bundled decoder_context struct for measurement extraction by @eliotheinrich in #671
  • Virtualize methods in realtime decoder API by @bmhowe23 in #674
  • Consolidate decoder RPC wire format into a single header by @bmhowe23 in #681
  • Move nv-qldpc-decoder schema registration into its plugin by @bmhowe23 in #701
  • Decoding server: add virtual hooks for decoder plugins to set D/O sparse matrices by @tlshannon in #746

Realtime infrastructure — fixes & test coverage (experimental):

nv-qldpc-decoder Updates (Closed Source)

New features and options

  • Gamma-ensemble sequential relay-BP (@kvmto) — a new sparse-GPU kernel adds a gamma_ensemble_size option (1/2/4/8; default 1 = disabled) that runs multiple parallel gamma "lanes" per relay iteration with race-to-fastest semantics (first lane to satisfy the stopping criterion wins; ties broken by lowest-weight correction). Supported on the sparse-GPU single-decode path with composition=1 and bp_method=3 or 5.
  • Sum-product BP variants (@bmhowe23) — bp_method gains 4 (sum-product + memory) and 5 (sum-product + damped memory), both requiring use_sparsity=True. Sequential relay (composition=1) now accepts bp_method=5, with gamma0, gamma_dist, and explicit_gammas extended to the new methods.
  • Native scipy.sparse parity-check-matrix input (@vedika-saravanan, @bmhowe23) — the decoder was migrated to the sparse_binary_matrix API and accepts any scipy.sparse format (CSR/CSC/COO/…) directly, with no dense .toarray()/.todense() conversion.
  • On-device observables output (@melody-ren) — a new optional O matrix (shape num_observables × block_size) makes decode()/decode_batch() return observable flips (O · correction mod 2) directly instead of the raw correction vector.

Correctness fixes

  • Fixed a race condition and undersized allocations in the batched and persistent-buffer sparse-GPU BP paths (@bmhowe23).
  • Fixed a sparse batched-GPU offset overflow and an out-of-bounds access in the OSD solver (osd_solver_gf2) (@melody-ren).

Build / packaging

  • AArch64 builds now target -march=armv8-a (was -march=native) for portability; x86-64 remains at x86-64-v3 (AVX2/SSE2).
  • Wired the plugin into the new declarative decoder-config schema.

Bug Fixes (QEC) 🐛

  • Fix non-physical TensorNetworkDecoder posteriors by @vedika-saravanan in #707
  • Fix PyMatching decoding for observable indices above 31 by @kaiqiy-nv in #697
  • Fix surface-code hook errors with an orientation-aware CNOT schedule by @bmhowe23 in #672
  • Fix Python QEC conversion for empty repetition-code X matrices by @kaiqiy-nv in #660
  • Stabilizer round kernels in Python no longer fail when the return-type annotation is incorrect by @eliotheinrich in #642
  • Fix OOB r/w when batch failure returns empty results by @kaiqiy-nv in #625
  • Fix DEM canonicalization to preserve the modeled error process by @bmhowe23 in #610
  • Throw an error if no noise is provided to DEM Python binding functions by @kaiqiy-nv in #565
  • Fix typo on pointer access by @kaiqiy-nv in #535
  • Fix potential use-after-free and add checks before heavy calculations by @kaiqiy-nv in #500

Breaking Changes & Deprecations ⚠️

  • Sparse parity-check matrix decoder API (#550): decoder constructors now take const sparse_binary_matrix &. Dense callers keep working via implicit conversion, but out-of-tree decoder plugins must be recompiled.
  • Detector-based memory circuits (#636): sample_memory_circuit syndrome output shape changes from (nShots·nRounds) × nAncilla to nShots × nDetectors (Stim-aligned). Adds x/z_sample_memory_circuit and a decompose_errors flag on dem_from_memory_circuit.
  • NumPy decode results (#558): DecoderResult.result (and decode_batch, async, and tuple-unpacking variants) now return a 1-D NumPy array instead of a Python list.
  • Pybind → Nanobind (#512): Python bindings migrated to Nanobind (for CUDA-Q compatibility) and the cuquantum-python dependency was bumped. A build/packaging change for downstream consumers.
  • Revert noise-learning support (#714): reverts the differentiable noise-learning (NMOptimizer) support for the TN decoder that was introduced in #526 and #577. Do not rely on that API in 0.7.0.
  • Remove hololink_gpu_id env var (#692): removed in favor of the cuda_device_id placement knob; reconcile_gpu_roce_device renamed to resolve_decode_device.
  • Remove --enable-mlir build flag (#586) and remove usage of CUDA-Q library mode (#523).
  • Rename CUDAQ_REGISTER_TYPE → CUDAQ_EXT_PT_REGISTER_TYPE (#542).

Documentation ✏️

Common / Misc

CUDA-Q decoupling & dependency cleanup

CUDA-Q version alignment (LLVM 22 / namespace unification / bumps)

CI / images / wheels / build plumbing

Tests / coverage