CUDA-Q QEC 0.8.0
This is a release of CUDA-Q QEC version 0.8.0. (CUDA-Q Solvers is not included in this release; see CUDA-Q Algorithms for a replacement for CUDA-Q Solvers.)
CUDA-Q QEC 0.8.0 is a decoder-focused release centered on new decoder options. The most notable change is:
- A new NV Fusion Decoder.
nv_fusion_decoderis a multi-threaded minimum-weight perfect matching decoder that combines fusion blossom with PyMatching's sparse blossom: the matching graph is partitioned into temporal blocks solved independently and then fused across their boundaries, dispatched to a worker pool as syndrome data arrives, so it is designed for low-latency streaming/realtime decoding as well as offline batch decoding (contributed by @tlshannon). This is currently closed source, so source code for this decoder is not available in GitHub at this time.
Additionally, nv-qldpc-decoder (closed source) gains two new features. The osd_init_method="min_llr" option makes BP+OSD far more accurate on full circuit-level (joint XYZ) decoding problems, cutting the logical error rate by up to 117x while running significantly fewer BP iterations. The relay_solutions helper records every Relay BP convergence during a decode, so the whole RelayBP-N stop_nconv trade-off curve can be computed offline from a single run instead of re-decoding once per N value (61 runs in the documented example).
Please check out the docs and examples for how to get started using CUDA-Q QEC.
Dependency note: CUDA-Q QEC 0.8.0 depends on CUDA-Q 0.16.
Features and Enhancements (QEC) 🎉
Decoders
- Add NV Fusion Decoder QEC plugin - a new multi-threaded fusion-blossom/sparse-blossom MWPM decoder for streaming and batch decoding, with decoder stats, latency speedups, and auto-selected
block_leaf_sizeby @tlshannon. - Allow a realtime decoder to be constructed from a raw Stim DEM (
stim_dem_pathin YAML) by @tlshannon in #801 — a third way (alongside flat-matrix anddem_chunksforms) to describe a decoder's error model on the realtime path, enabling DEM-native decoders such as Chromobius there. In particular, this lets a realtime color-code predecoder pipeline chain Chromobius in as the global decoder (stim_dem_pathreaches the nested global decoder viaglobal_decoder_params["stim_dem"]), since Chromobius rejects a parity-check-matrix representation outright and previously could not be configured on the realtime path at all. See NVIDIA/Ising-Decoding for the color-code predecoder this is designed to pair with. - Add common decoder-stats functions (
decoder_stats) for consistent latency/replay logging across decoders by @tlshannon in #802. - Add
[Core]rvalueinsert(key, T&&)overload toheterogeneous_map, avoiding unnecessary deep copies of large payloads (e.g. syndrome/LLR history) by @bmhowe23 in #781.
Bindings / examples
- Replace
msmexecution contexts withcudaq::dem_from_kernelin app examples by @bmhowe23 in #696. - Accept both
cc.StdvecTypeandcc.SequenceTypeinpy_code, tracking a CUDA-Q internal rename by @bmhowe23 in #786.
nv-qldpc-decoder Updates (Closed Source)
New features and options
osd_init_method="min_llr"(@bmhowe23) — a new OSD seeding mode that initializes the OSD solver with the minimum-LLR-magnitude scoreboard rather than the existing default. This makes BP+OSD far more accurate on full circuit-level (joint XYZ) decoding problems while running significantly fewer BP iterations.relay_solutionspost-processing helper (@bmhowe23) — a pure-NumPy module that reconstructs, offline, what relay-BP would have returned for anystop_nconvsetting from a single uncapped recording run, plusbatch_opt_resultsplumbing to carryrelay_solutionsrecords through the decoder API. It records every relay-BP convergence (not just the accepted one).- Speed up BP LLR-history collection on the sparse-GPU path (@bmhowe23).
Bug Fixes (QEC) 🐛
- Fix PyMatching parallel-edge merge mapping and realtime observable output (native
decode_to_obsoutput instead of an error-frame-to-observable projection) — cuts realtime PyMatching decode latency by ~30-35% on GB200 by @vedika-saravanan in #799. - Reduce realtime decode latency by resolving detectors as measurements arrive and reusing PyMatching scratch buffers (steady-state
decode()now allocates nothing) by @tlshannon in #812. - Fix decoder lifetime during async decoding by @kaiqiy-nv in #711.
- Validate AI decoder RPC slot capacity to prevent gateway output from overwriting adjacent slots by @kaiqiy-nv in #764.
Breaking Changes & Deprecations ⚠️
- CUDA-Q Solvers removed from this repository (#796, #797): Solvers is deprecated in favor of the new CUDA-Q Algorithms library; future development in this repo targets QEC only, and Algorithms development continues at https://github.com/NVIDIA/cudaq-algorithms.
Documentation ✏️
- Add a gamma-ensemble Relay-BP performance-tuning guide by @eliotheinrich in #755, plus the accompanying standalone benchmarking scripts by @eliotheinrich in #766.
- Correct realtime decoding example commands (
surface_code_1.pyfilename,--emulateopt-in, Quantinuum payload provider args) by @vedika-saravanan in #749. - Add versioning to the docs site (
gh-pagesversion switcher) by @anjbur in #761. - Update doc layout to match the 0.7.0 release branch's layout while keeping main's content by @melody-ren in #760.
Common / Misc
CI / build plumbing
- Give each
build_wheels.yamlinstance a distinct artifact name by @bmhowe23 in #767. - Update validation-wheels script for CUDA/torch package-compatibility issues and ARM TensorRT availability by @kaiqiy-nv in #817.
- Update sync workflow to fetch LFS objects from upstream (fixing 404s when upstream adds new LFS files) and to use unauthenticated access to the public repo by @anjbur in #757 / #762.
- Make the CUDA-Q build stage of the dev image optional (parametrized
FROM) so a realtime-development image can be composed from cudaqx + CUDA-Q realtime docker assets, default behavior unchanged, by @Renaud-K in #793.
Release housekeeping