Releases: NVIDIA/cudaq-qec
Release list
0.8.0
CUDA-Q QEC 0.8.0
This is a release of CUDA-Q QEC version 0.8.0. (CUDA-Q Solvers is not included in this release; see CUDA-Q Algorithms for a replacement for CUDA-Q Solvers.)
CUDA-Q QEC 0.8.0 is a decoder-focused release centered on new decoder options. The most notable change is:
- A new NV Fusion Decoder.
nv_fusion_decoderis a multi-threaded minimum-weight perfect matching decoder that combines fusion blossom with PyMatching's sparse blossom: the matching graph is partitioned into temporal blocks solved independently and then fused across their boundaries, dispatched to a worker pool as syndrome data arrives, so it is designed for low-latency streaming/realtime decoding as well as offline batch decoding (contributed by @tlshannon). This is currently closed source, so source code for this decoder is not available in GitHub at this time.
Additionally, nv-qldpc-decoder (closed source) gains two new features. The osd_init_method="min_llr" option makes BP+OSD far more accurate on full circuit-level (joint XYZ) decoding problems, cutting the logical error rate by up to 117x while running significantly fewer BP iterations. The relay_solutions helper records every Relay BP convergence during a decode, so the whole RelayBP-N stop_nconv trade-off curve can be computed offline from a single run instead of re-decoding once per N value (61 runs in the documented example).
Please check out the docs and examples for how to get started using CUDA-Q QEC.
Dependency note: CUDA-Q QEC 0.8.0 depends on CUDA-Q 0.16.
Features and Enhancements (QEC) 🎉
Decoders
- Add NV Fusion Decoder QEC plugin - a new multi-threaded fusion-blossom/sparse-blossom MWPM decoder for streaming and batch decoding, with decoder stats, latency speedups, and auto-selected
block_leaf_sizeby @tlshannon. - Allow a realtime decoder to be constructed from a raw Stim DEM (
stim_dem_pathin YAML) by @tlshannon in #801 — a third way (alongside flat-matrix anddem_chunksforms) to describe a decoder's error model on the realtime path, enabling DEM-native decoders such as Chromobius there. In particular, this lets a realtime color-code predecoder pipeline chain Chromobius in as the global decoder (stim_dem_pathreaches the nested global decoder viaglobal_decoder_params["stim_dem"]), since Chromobius rejects a parity-check-matrix representation outright and previously could not be configured on the realtime path at all. See NVIDIA/Ising-Decoding for the color-code predecoder this is designed to pair with. - Add common decoder-stats functions (
decoder_stats) for consistent latency/replay logging across decoders by @tlshannon in #802. - Add
[Core]rvalueinsert(key, T&&)overload toheterogeneous_map, avoiding unnecessary deep copies of large payloads (e.g. syndrome/LLR history) by @bmhowe23 in #781.
Bindings / examples
- Replace
msmexecution contexts withcudaq::dem_from_kernelin app examples by @bmhowe23 in #696. - Accept both
cc.StdvecTypeandcc.SequenceTypeinpy_code, tracking a CUDA-Q internal rename by @bmhowe23 in #786.
nv-qldpc-decoder Updates (Closed Source)
New features and options
osd_init_method="min_llr"(@bmhowe23) — a new OSD seeding mode that initializes the OSD solver with the minimum-LLR-magnitude scoreboard rather than the existing default. This makes BP+OSD far more accurate on full circuit-level (joint XYZ) decoding problems while running significantly fewer BP iterations.relay_solutionspost-processing helper (@bmhowe23) — a pure-NumPy module that reconstructs, offline, what relay-BP would have returned for anystop_nconvsetting from a single uncapped recording run, plusbatch_opt_resultsplumbing to carryrelay_solutionsrecords through the decoder API. It records every relay-BP convergence (not just the accepted one).- Speed up BP LLR-history collection on the sparse-GPU path (@bmhowe23).
Bug Fixes (QEC) 🐛
- Fix PyMatching parallel-edge merge mapping and realtime observable output (native
decode_to_obsoutput instead of an error-frame-to-observable projection) — cuts realtime PyMatching decode latency by ~30-35% on GB200 by @vedika-saravanan in #799. - Reduce realtime decode latency by resolving detectors as measurements arrive and reusing PyMatching scratch buffers (steady-state
decode()now allocates nothing) by @tlshannon in #812. - Fix decoder lifetime during async decoding by @kaiqiy-nv in #711.
- Validate AI decoder RPC slot capacity to prevent gateway output from overwriting adjacent slots by @kaiqiy-nv in #764.
Breaking Changes & Deprecations ⚠️
- CUDA-Q Solvers removed from this repository (#796, #797): Solvers is deprecated in favor of the new CUDA-Q Algorithms library; future development in this repo targets QEC only, and Algorithms development continues at https://github.com/NVIDIA/cudaq-algorithms.
Documentation ✏️
- Add a gamma-ensemble Relay-BP performance-tuning guide by @eliotheinrich in #755, plus the accompanying standalone benchmarking scripts by @eliotheinrich in #766.
- Correct realtime decoding example commands (
surface_code_1.pyfilename,--emulateopt-in, Quantinuum payload provider args) by @vedika-saravanan in #749. - Add versioning to the docs site (
gh-pagesversion switcher) by @anjbur in #761. - Update doc layout to match the 0.7.0 release branch's layout while keeping main's content by @melody-ren in #760.
Common / Misc
CI / build plumbing
- Give each
build_wheels.yamlinstance a distinct artifact name by @bmhowe23 in #767. - Update validation-wheels script for CUDA/torch package-compatibility issues and ARM TensorRT availability by @kaiqiy-nv in #817.
- Update sync workflow to fetch LFS objects from upstream (fixing 404s when upstream adds new LFS files) and to use unauthenticated access to the public repo by @anjbur in #757 / #762.
- Make the CUDA-Q build stage of the dev image optional (parametrized
FROM) so a realtime-development image can be composed from cudaqx + CUDA-Q realtime docker assets, default behavior unchanged, by @Renaud-K in #793.
Release housekeeping
0.7.0
CUDA-Q QEC 0.7.0
This is a release of CUDA-Q QEC version 0.7.0. (CUDA-Q Solvers is not included in this release.)
CUDA-Q QEC 0.7.0 is a decoder-focused release that makes the decoder stack faster, more scalable under real-time constraints, and easier to extend. The most notable changes are:
- Lower logical error rates under real-time deadlines. The closed-source
nv-qldpc-decoderadds a new gamma-ensemble Relay-BP mode (contributed by @kvmto) that runs several independent gamma trajectories ("lanes") in parallel on a single GPU and exits as soon as any lane converges. By contracting the decoder's latency tail — what matters most when every decode must finish within a fixed wall-clock budget, as in real-time QEC — it lowers the logical error rate under fixed decoding deadlines by up to ~89× for large bivariate-bicycle qLDPC codes on a GB200 (for example ~89× for[[288,12,18]]and ~41× for[[144,12,12]]at deadlines of ~1–5 ms), while also tightening worst-case (p99.99) latency by ~2.7-5.5×. See Improving Relay BP Decoding With Gamma Ensembles for more details. - Sparse parity-check matrices end to end, so large qLDPC codes no longer need their parity-check matrices materialized as dense tensors — including native
scipy.sparseinput for thenv-qldpc-decoder. - Decoder construction directly from Stim detector-error-model (DEM) strings, plus a new DEM-native Chromobius color-code decoder.
- A GPU/CPU
dem_samplingcapability - Surface-code orientation (XV/XH/ZV/ZH) control.
- A declarative decoder-configuration schema that lets third-party decoders be fully YAML-configurable from their own shared library with no CUDA-Q QEC rebuild.
Beyond gamma-ensemble, the nv-qldpc-decoder also gains sum-product BP variants and on-device observables output, along with several performance and correctness fixes. Under the hood, the Python bindings were migrated from Pybind11 to Nanobind for upstream CUDA-Q compatibility, and the decoder/realtime code paths were decoupled from the CUDA-Q runtime (including a dedicated CUDA-Q QEC logger).
Please check out the docs and examples for how to get started using CUDA-Q QEC.
Dependency note: CUDA-Q QEC 0.7.0 depends on CUDA-Q 0.15.1 and builds against its published images (CUDA 12.6 and 13.0).
Realtime decoding: The real-time GPU decoding capabilities introduced in 0.6.0 (the CUDA-Q Realtime
HOST_LOOPbridge and the Relay-BP / PyMatching predecoder examples) continue to work in 0.7.0.
Features and Enhancements (QEC) 🎉
Decoders & detector error models
- Sparse parity-check matrix support for decoders by @vedika-saravanan in #550
- Sparse-aware PCM utility migration by @vedika-saravanan in #602
- Make canonicalization a member of
sparse_binary_matrixby @vedika-saravanan in #599 - Adopt
scipy.sparseas optional interop by @bmhowe23 in #590 - Fix dense → sparse conversion in
get_decoderto avoid redundant copies by @bmhowe23 in #589 - Support Stim DEM strings in
get_decoderby @vedika-saravanan in #571 - DEM from Stim text can use error decompositions by @eliotheinrich in #615
- Add
dem_samplingwith CPU and GPU backends (C++ and Python) by @kvmto in #479 - Add Chromobius decoder to the decoder plugins by @wsttiger in #546
- Add YAML and Python config support for Chromobius / TRT global decoder by @melody-ren in #633
- Extend
trt_decoderwith global decoder chaining by @bmhowe23 in #524 - Expand TRT decoder YAML config for composite decoding by @wsttiger in #536
- Add CLI override flags and
from_nametoPipelineConfigby @wsttiger in #503 - Fix
ai_decoder_serviceTRT builder for quantized ONNX (FP8) by @wsttiger in #507 - Add boundary-aware overloads for
canonicalize_for_roundsand sliding-window decoder by @eliotheinrich in #656 - Use canonical soft-to-hard decoder conversion by @melody-ren in #553
- Return Python decode result as NumPy arrays by @melody-ren in #558
trt_decodernow throws on inference failure instead of returning stale/zeroed results by @melody-ren in #680
Note (composite decoding): The
trt_decodercan chain a second-stage "global" decoder (for example PyMatching or Chromobius) via theglobal_decoderandglobal_decoder_paramsoptions. When constructing the decoder directly (rather than from a YAML config, which fills this in automatically), you must supplyglobal_decoder_params— an empty map is fine — wheneverglobal_decoderis set, or the global stage is skipped.
Codes & circuits
- Add surface-code orientation support (XV/XH/ZV/ZH) by @melody-ren in #637
- Add detector support to memory circuits by @eliotheinrich in #636
Extensibility & diagnostics
- Declarative decoder parameter schemas: pluggable realtime decoder configuration by @bmhowe23 in #679
- Create CUDA-Q QEC logger by @tlshannon in #630
- Add
cuda_device_idplacement knob for GPU decoders by @melody-ren in #690 - Route every decoder device pin through one resolver and two wrappers by @melody-ren in #698
Realtime decoding infrastructure (experimental)
- Standalone realtime QEC decoding server by @bmhowe23 in #666
- Add QEC decoder-server core and CQR adapter by @vedika-saravanan in #653
- Add host-side in-process-RPC path for real-time QEC decoding by @cketcham2333 in #609
- Add PyMatching HOST_CALL decoder-server RPC path by @cketcham2333 in #600
- Add vanilla PyMatching support to realtime decoder config by @vedika-saravanan in #614
- qec/realtime: device-graph scheduler for per-round Hololink QLDPC decoding by @cketcham2333 in #631
- Add bundled
decoder_contextstruct for measurement extraction by @eliotheinrich in #671 - Virtualize methods in realtime decoder API by @bmhowe23 in #674
- Consolidate decoder RPC wire format into a single header by @bmhowe23 in #681
- Move nv-qldpc-decoder schema registration into its plugin by @bmhowe23 in #701
- Decoding server: add virtual hooks for decoder plugins to set
D/Osparse matrices by @tlshannon in #746
Realtime infrastructure — fixes & test coverage (experimental):
- Fix GB200
gpu_rocedecoding-server validation issues by @vedika-saravanan in #683 - Add QEC
device_callhost-dispatch E2E coverage by @vedika-saravanan in #628 - Relay BP (nv-qldpc)
gpu_roceprofile for the HSB decoding-server test by @cketcham2333 in #670 - Add a TRT + PyMatching decoder profile to the HSB FPGA decoding-server test by @cketcham2333 in #673
- Fix CQR surface-code test to use Relay-BP by @vedika-saravanan in #691
nv-qldpc-decoder Updates (Closed Source)
New features and options
- Gamma-ensemble sequential relay-BP (@kvmto) — a new sparse-GPU kernel adds a
gamma_ensemble_sizeoption (1/2/4/8; default 1 = disabled) that runs multiple parallel gamma "lanes" per relay iteration with race-to-fastest semantics (first lane to satisfy the stopping criterion wins; ties broken by lowest-weight correction). Supported on the sparse-GPU single-decode path withcomposition=1andbp_method=3or5. - Sum-product BP variants (@bmhowe23) —
bp_methodgains4(sum-product + memory) and5(sum-product + damped memory), both requiringuse_sparsity=True. Sequential relay (composition=1) now acceptsbp_method=5, withgamma0,gamma_dist, andexplicit_gammasextended to the new methods. - Native
scipy.sparseparity-check-matrix input (@vedika-saravanan, @bmhowe23) — the decoder was migrated to thesparse_binary_matrixAPI and accepts anyscipy.sparseformat (CSR/CSC/COO/…) directly, with no dense.toarray()/.todense()conversion. - On-device observables output (@melody-ren) — a new optional
Omatrix (shapenum_observables × block_size) makesdecode()/decode_batch()return observable flips (O · correction mod 2) directly instead of the raw correction vector.
Correctness fixes
- Fixed a race condition and undersized allocations in the batched and persistent-buffer sparse-GPU BP paths (@bmhowe23).
- Fixed a sparse batched-GPU offset overflow and an out-of-bounds access in the OSD solver (
osd_solver_gf2) (@melody-ren).
Build / packaging
- AArch64 builds now target
-march=armv8-a(was-march=native) for portability; x86-64 remains at x86-64-v3 (AVX2/SSE2). - Wired the plugin into the new de...
0.6.0
CUDA-Q QEC 0.6.0 and CUDA-Q Solvers 0.6.0
This is combined release of CUDA-Q QEC and CUDA-Q Solvers, both version 0.6.0.
This is the first CUDA-Q QEC release that builds example decoder applications on top of CUDA-Q Realtime [blog]. CUDA-Q QEC 0.6 ships with two new real-time-capable decoder pipelines: the RelayBP belief-propagation decoder for qLDPC codes and an NVIDIA Ising convolutional neural network (CNN) pre-decoder paired with a global decoder (PyMatching) for the surface code. These pipelines enable quantum vendors and QEC researchers to deploy real-time GPU decoding for two popular code families via NVQLink.
Additionally, this release of CUDA-Q QEC contains speed improvements for our GPU-accelerated RelayBP decoder (up to 19X!)
For CUDA-Q Solvers 0.6.0, support was added for a new UpCCGSD ansatz solver and a Coupled Exchange Operator (CEO) pool.
Please check out the docs and examples for how to get started using the CUDA-QX libraries!
Note: CUDA-Q QEC 0.6.0 and CUDA-Q Solvers 0.6.0 both depend on CUDA-Q 0.14. For CUDA-Q Realtime usage (experimental), you need to use CUDA-Q 0.14.1.
Features and Enhancements (QEC) 🎉
- Sliding window optimize by @cketcham2333 in #343
- feat(qec): add trt_decoder_config for real-time decoding by @wsttiger in #384
- Add trt cudagraph by @wsttiger in #369
- Add decode batch by @wsttiger in #383
- Create PyMatching decoder plugin by @bmhowe23 in #396
- Add optional
Oparameter to PyMatching plugin by @bmhowe23 in #449 - Update
trt_decoderto supportuint8data types for I/O by @bmhowe23 in #455 - Add graph capture functions to common decoder interface by @bmhowe23 in #475
- Follow-up to #475 - additional decoder interface updates by @bmhowe23 in #478
- Add Hololink QLDPC graph decode bridge and CI test by @cketcham2333 in #481
- Add realtime AI decoder / predecoder infrastructure (GPU + Host) w/ host dispatcher by @wsttiger in #457
- Add FPGA-based test application for realtime predecoder by @wsttiger in #490
nv-qldpc-decoder Updates (Closed Source)
-
Implemented new graph capture interface functions to run RelayBP with CUDA-Q Realtime
-
Add new
repeatableconfiguration option to enable bit-for-bit repeatable results when running back-to-back on the same system -
Significant RelayBP optimizations, for both
fp32andfp64. Timings below show speedups relative to 0.5 for some well-known Bicycle Bivariate Codes on B200. All of the reported speedups are for non-batched, serial execution mode.Case Name (n_k_d) Variant Total Speedup (Ratio) 72_12_6 fp32 3.44 72_12_6 fp64 2.74 144_12_12 fp32 5.50 144_12_12 fp64 4.51 288_12_18 fp32 19.06 288_12_18 fp64 13.16 Average 8.07
Bug Fixes (QEC) 🐛
- Coverity fixes by @cketcham2333 in #395
- Fix OOB r/w issues by @kaiqiy-nv in #464
- Add onnxscript to trt_decoder optional dependency by @wsttiger in #407
- Bug fix and add test cases by @kaiqiy-nv in #424
- Fix pytorch AcceleratorError root-caused by QEC by @kaiqiy-nv in #472
Features and Enhancements (Solvers) 🎉
- UpCCGSD ansatz solver by @rr637 in #372
- Add Coupled Exchange Operator (CEO) pool by @jgonthier in #387
Bug Fixes (Solvers) 🐛
- Fix OOB r/w issues by @kaiqiy-nv in #464
- Fix the gradient evaluation bugs and add test cases by @kaiqiy-nv in #434
- Fix optimiser forwarding bug and other minor bugs and add test cases by @kaiqiy-nv in #435
- Fix BK transformation | Fix JW parity Z-chain | Add test cases by @kaiqiy-nv in #460
- Fix GQE invalid CUDA handle / AcceleratorError when moving model to GPU by @vedika-saravanan in #473
- Fix mixer forwarding | Fix MPI implementation | Add test cases by @kaiqiy-nv in #452
- Add error message for missing system dependencies by @vedika-saravanan in #467
- Fix qubit indices bug and add test case by @kaiqiy-nv in #469
- GQE: PyTorch GPU compatibility check, exit/skip test on mismatch by @vedika-saravanan in #494
Documentation ✏️
- Update documentation for relay-bp by @melody-ren in #346
- [Docs] Add uccgsd to doc by @marwafar in #340
- [Docs] update gen_ham with UHF by @marwafar in #339
- [docs] Update nv-qldpc-decoder docs to describe the new proc_float option by @bmhowe23 in #288
- Incorporate Sliding Window Decoder docs by @bmhowe23 in #359
- Add docs for realtime decoding by @kvmto in #345
- Add docs for AI decoder training with PyTorch by @wsttiger in #344
- fix typo in calling operator pool with uccsd in doc by @marwafar in #366
- Add requirement for memory BP methods in docs by @melody-ren in #376
- Added trt_decoder docs for Python and C++ by @wsttiger in #381
Common / Misc
- License updates by @bmhowe23 in #360
- Add license agreement notification to Docker image by @bmhowe23 in #362
- [core] Fix pre-existing extension point issue by @bmhowe23 in #374
- Bump CUDA-Q commit (with support for breaking changes) by @github-actions[bot] in #416
- Bump CUDA-Q commit (with non-trivial updates) by @github-actions[bot] in #437
- Bump CUDA-Q commit and re-enable some tests by @bmhowe23 in #450
- Bump CUDA-Q dependencies from 0.13 to 0.14 by @bmhowe23 in #468
- Align CUDA-Q and CUDA-Q Realtime commits for 0.14.1 by @bmhowe23 in #489
- Fix minor issues reported by Coverity by @kaiqiy-nv in #399
- Redundantly including Logger.h and FmtCore.h includes ahead of runtime refactor. by @Renaud-K in #409
- Update cuda-quantum-devdeps:ext-... to cuda-quantum-devcontainer-... by @bmhowe23 in #420
- Fix build if CUDAQ_REALTIME_ROOT is not set by @bmhowe23 in #432
- Follow-up to #396 and #416 - fix wheel builds by @bmhowe23 in #439
- Update heterogeneous_map to recognize ints as bools by @bmhowe23 in #441
- Update CMake for TensorRT decoder unit test by @bmhowe23 in #448
Testing
- Update test scripts to allow easy data generation by @bmhowe23 in #354
- Follow-up to #344 - add onnxscript to wheels test env by @bmhowe23 in #368
- [uccsd/uccgsd] Update test tolerances by @bmhowe23 in #375
- Add playback/record to the surface code 1 test by @cketcham2333 in #406
- Add nv-qldpc-test option to surface_code-1.cpp by @cketcham2333 in #415
- Mock decoder by @cketcham2333 in #423
- Add cuda graph launch to mock decoder and introduce autonomous decoder CRTP by @cketcham2333 in #429
- [ci] Update how to get cudaq::realtime by @bmhowe23 in #443
- [realtime] Mock decoder updates for latest cudaq::realtime updates by @bmhowe23 in #444
- [realtime] Move mock decoders to test directories by @bmhowe23 in #445
- [realtime] Advance cudaq::realtime commit and update names accordingly by @bmhowe23 in #446
- Skip GQE GPU tests on unsupported GPU architectures by @vedika-saravanan in #477
- Update container validation script by @bmhowe23 in #487
- Remove cudaq-realtime from All libs CI and enable QLDPC graph test in Release CI by @cketcham2333 in #482
New Contributors
- @kaiqiy-nv made their first contribution in https://github.com/NVIDIA/cu...
0.5.0
CUDA-Q QEC 0.5.0 and CUDA-Q Solvers 0.5.0
This is combined release of CUDA-Q QEC and CUDA-Q Solvers, both version 0.5.0.
This release introduces three major new decoders - our TensorRT AI/ML decoder, a GPU-accelerated implementation of the RelayBP decoding algorithm [1], and a new Sliding Window decoder.
CUDA-Q QEC 0.5.0 also includes our first real-time decoder API, enabling true in-kernel decoding for quantum error-correcting codes implemented directly in CUDA-Q! This makes it possible to integrate real-time QEC decoding into device-side kernels and will allow richer experimentation both in simulation and on real hardware [2].
Additionally, support for CUDA 13 is added in this release. Many more features are listed below!
Note:
Support for Python 3.10 has been removed in this release.
Please check out the docs and examples for how to get started using the CUDA-QX libraries!
Note: CUDA-Q QEC 0.5.0 and CUDA-Q Solvers 0.5.0 both depend on CUDA-Q 0.13.0.
[1] https://arxiv.org/abs/2506.01779
[2] https://docs.quantinuum.com/systems/trainings/helios/getting_started/gpu_decoding.html
Features and Enhancements (QEC) 🎉
- Generalize
single_error_luttomulti_error_lutby @bmhowe23 in #281 - Create sliding window decoder by @bmhowe23 in #198
- Add TensorRT decoder by @wsttiger in #307
- Real-time decoding support by @bmhowe23 in #333
Bug Fixes (QEC) 🐛
- Fix issue #258: added check and informative error message by @kvmto in #261
- Fix tensor_network_decoder.py warning by @bmhowe23 in #328
nv-qldpc-decoder Updates (Closed Source)
- Update the
nv-qldpc-decoderto support RelayBP by adding new options into the existing framework - Update the
nv-qldpc-decoderto support aproc_floatoption to select fp32 processing (instead of fp64 default)
Features and Enhancements (Solvers) 🎉
Bug Fixes (Solvers) 🐛
Documentation ✏️
- Add workaround for Blackwell by @melody-ren in #247
- Doc: update the supported gpu and python versions by @caiyunh in #251
- Add API doc update for BP by @melody-ren in #243
- Update gqe example for mpi by @melody-ren in #245
- Add instructions for installing dependencies of optional python components by @melody-ren in #274
- Fix rendering of doc by @melody-ren in #276
- Updating qec docs by @sacpis in #283
- Add CMake options description in Building.md by @melody-ren in #285
- Fix invalid reference to cudaqx-config on QEC Introduction docs page by @bmhowe23 in #287
- Improve capitalization and add developer information of PRIMA by @zaikunzhang in #290
- Add libgfortran to NOTICE file by @bmhowe23 in #292
- Fix missing docstring for solvers lib by @melody-ren in #297
- Added python test procedure by @cketcham2333 in #322
Common / Misc
- Incorporate validation script updates from releases/v0.4.0 branch by @bmhowe23 in #253
- Delay target specification until execution by @bmhowe23 in #244
- Add patch_wheel_metadata.sh script by @bmhowe23 in #275
- Add prune_cudaqx-dev_by_sha.sh helper script by @bmhowe23 in #280
- Update devdeps image to Ubuntu 24.04 by @bmhowe23 in #302
- Fix dev images (CUDA 12.6 instead of 12.0) by @bmhowe23 in #306
- Remove Python 3.10 and update test images by @bmhowe23 in #305
- [ci] Use Python 3.12 for yapf by @bmhowe23 in #310
- Advance from python3.10 to 3.11+ by @melody-ren in #312
- Prepare for CUDA 13 by @mitchdz in #308
- bump cuquantum 25.06 -> 25.09, bump cudaq commit by @mitchdz in #318
- Add CUDA 13 support and add metapackages by @bmhowe23 in #320
- Fix address sanitizer issue by @bmhowe23 in #323
- Fixed double free error by @cketcham2333 in #327
- Misc fixes for various build/test environments by @bmhowe23 in #331
- Fix TRT Decoder CMake for unsupported platforms by @bmhowe23 in #332
- Removed num_syndromes_per_round by @cketcham2333 in #338
- Cleanup the extra payload provider support by @bmhowe23 in #342
- Script and config updates in preparation for 0.5.0 by @bmhowe23 in #341
- Torch/TensorRT wheel compatibility updates by @bmhowe23 in #350
- Don't throw error if optional TRT dependencies are not installed by @bmhowe23 in #352
- Bump cuQuantum 25.09 -> 25.09.1 by @bmhowe23 in #351
Testing
- Add tests for qec lib by @caiyunh in #226
- Test python examples in build_wheels job by @melody-ren in #256
- Use 'tensor-network-decoder' in pip install instructions by @melody-ren in #263
- Change v100 runners to a100 by @bmhowe23 in #266
- Update Tensor Network Decoder docs to address cuTensor 2.3 by @bmhowe23 in #268
- Follow-up to #268 - fix link rendering by @bmhowe23 in #269
- [docs] Update Building from Source section by @bmhowe23 in #270
- Bug 5652917: update scripts for the failed test gqe_h2.py by @caiyunh in #353
New Contributors
- @sacpis made their first contribution in #283
- @zaikunzhang made their first contribution in #290
- @cketcham2333 made their first contribution in #322
- @marwafar made their first contribution in #309
Full Changelog: 0.4.0...0.5.0
0.4.0
The CUDA-QX 0.4.0 release includes a variety of new features in both the Solvers and QEC library. For the Solvers library, a new Generative Quantum Eigensolver implementation is provided. To use this new algorithm, you will need to use pip install cudaq-solvers[gqe] in order to install all the proper dependencies. For the QEC library, a new Tensor Network Decoder is added, and a new API allows users to automatically generate PCMs from noisy CUDA-Q memory circuits. To use this new decoder, you will need to a) use Python >= 3.11, and b) use pip install cudaq-qec[tensor_network_decoder] in order to install all the proper dependencies. Additionally, support for Python 3.13 is added in this release. Many more features are listed below!
Please check out the docs and examples for how to get started using the CUDA-QX libraries!
Note: CUDA-QX 0.4.0 depends on CUDA-Q 0.12.0.
Features and Enhancements (Solvers) 🎉
- Generative Quantum Eigensolver by @melody-ren in #201
Features and Enhancements (QEC) 🎉
- Tensor Network Decoder by @npancotti in #179
- Add check for tn decoder import by @melody-ren in #206
- Add Parity Check Matrix generation and utility functions by @bmhowe23 in #147
- Add
opt_resultsfield todecoder_resultby @melody-ren in #171 - Optimize
convert_vec_soft_to_tensor_hardby 50x by @bmhowe23 in #160
Breaking Changes (QEC) 🛠
Bug Fixes (QEC) 🐛
nv-qldpc-decoder Updates (Closed Source)
- Added multiple algorithm configuration updates, including
iter_per_check,clip_value,bp_method(now supporting both sum-product and min-sum),scale_factor, and an optional output result logging capability forbp_llr_history.- Add API doc update for BP by @melody-ren #243
Documentation ✏️
- Replace PCM generation example with one using new DEM API by @kvmto in #216
- introduction.rst: clean up docs typos by @mitchdz in #163
- Update doc and example to use decoder_result new API by @melody-ren in #175
- Update examples to show how to use decoder_result API by @melody-ren in #184
- [docs] Update circuit-level noise example to use new DEM feature by @bmhowe23 in #199
- Fix typo in Building.md by @melody-ren in #209
- Doc update by @melody-ren in #213
- Tensor network decoder docs & examples by @npancotti in #214
- Update Python READMEs to describe optional features by @bmhowe23 in #236
- small fix in
gqedocstring by @caldwellshane in #239 - Update docs and guard for tensor network decoder by @melody-ren in #241
Common / Misc
- Add python 3.13 support by @melody-ren in #222
- [nfc] Be more explicit about extension_point usage by @bmhowe23 in #177
- Add move constructor to
heterogeneous_mapby @melody-ren in #186 - Remove unnecessary
MeasureCounts.hinclude by @1tnguyen in #188 - Update restrict pointer syntax to allow compilation with clang++ by @bmhowe23 in #205
- Update cudaq ver for project.toml by @melody-ren in #219
- Fix decoder plugin not throwing the missing dependecy err msg issue by @melody-ren in #224
- Update pyproject.toml dependencies for CUDA-Q 0.12 by @bmhowe23 in #228
- Add missing licence header by @melody-ren in #235
Testing
- Workflow updates (incl creating all_libs_release.yaml) by @bmhowe23 in #159
- [test] Skip nv-qldpc-decoder test if no GPUs found by @bmhowe23 in #238
- Add python3.13 to validate_wheels.sh by @melody-ren in #237
New Contributors
- @mitchdz made their first contribution in #163
- @1tnguyen made their first contribution in #188
- @npancotti made their first contribution in #179
Full Changelog: 0.3.0...0.4.0
0.3.0
The CUDA-QX 0.3.0 release includes updates for CUDA-Q breaking changes (spin ops) and notable performance improvements to our nv-qldpc-decoder. Some of these improvements reduce our average OSD-0 processing by >6X, so try it out!
For details on the breaking changes, see the descriptions in in PRs below along with the associated CUDA-Q PRs:
Please check out the docs and examples for how to get started using the CUDA-QX libraries!
What's Changed (Solvers)
- UCCSD operator pool correction and integration in adapt simulator by @kvmto in #65
- Update libraries for
cudaq::spin_opbreaking changes by @bmhowe23 in #134 - Update libraries for Python spin op breaking changes by @bmhowe23 in #152
- Add
__version__info to packages by @bmhowe23 in #154
What's Changed (QEC)
- Update libraries for
cudaq::spin_opbreaking changes by @bmhowe23 in #134 - Update libraries for Python spin op breaking changes by @bmhowe23 in #152
- Correcting syndromes per shot in circuit-level example by @justinlietz in #136
- Add
__version__info to packages by @bmhowe23 in #154
nv-qldpc-decoder Updates (Closed Source)
- OSD solver ARM performance improvements by @melody-ren
- Add Gaussian elimination to OSD with early exit @melody-ren
- Fix numerical stability bug with BP by @bmhowe23
Documentation
Testing
New Contributors
Note: CUDA-QX 0.3.0 depends on CUDA-Q 0.11.0.
Full Changelog: 0.2.1...0.3.0
0.2.1
What's Changed
This a is a minor patch release to fix 1 important bug and add 1 important set of error checks. Both of these changes are related the QEC decoders.
- Fix bug in nv-qldpc-decoder where exhaustive search OSD was not searching over the correct candidate bitstrings.
- Add error checking before constructing internal structures from Python data by @bmhowe23 #141
Note: CUDA-QX 0.2.1 depends on CUDA-Q 0.10.0 (same as CUDA-QX 0.2.0).
Full Changelog: 0.2.0...0.2.1
0.2.0
This release of the CUDA-QX libraries adds support for arm64 / aarch64 platforms for both the QEC and Solvers libraries.
CUDA-QX is a collection of libraries that build upon the CUDA-Q programming model to enable the rapid development of hybrid quantum-classical application code leveraging state-of-the-art CPUs, GPUs, and QPUs. It provides a collection of C++ libraries and Python packages that enable research, development, and application creation for use cases in quantum error correction and hybrid quantum-classical solvers.
Please check out the docs and examples for how to get started using the CUDA-QX libraries!
Note: CUDA-QX 0.2.0 depends on CUDA-Q 0.10.0.
What's Changed for QEC
This release includes a new high-performance GPU-accelerated QLDPC decoder implementation based on the algorithms described in Decoding Across the Quantum LDPC Code Landscape. This new decoder requires an NVIDIA GPU. Additionally, this release includes performance improvements in the sample_memory_circuit by utilizing CUDA-Q's new "explicit measurements" feature to accelerate collection of noisy syndrome data when using the stim target.
Features and Enhancements 🎉
- Introduce the new
nv-qldpc-decoder, a GPU-accelerated decoder of the Belief Propagation and Ordered Statistics Decoding algorithms described in https://arxiv.org/abs/2005.07016. - Make DecoderResult tuple-like in python by @melody-ren in #16
- Add Surface Code by @justinlietz in #7
- First shot memory by @justinlietz in #61
- Add python binding for decoder APIs by @melody-ren in #68
- Use explicit measurements for memory circuit experiments by @bmhowe23 in #89
- Support vector types for kwargs by @bmhowe23 in #99
Bug Fixes 🐛
Breaking Changes 🛠
- Change default decoder float type to double by @bmhowe23 in #31
- Update API name:
decode_multitodecode_batchby @melody-ren in #102
Documentation Updates ✏️
- Update docs to clarify Python wheels installation requirements by @bmhowe23 in #34
- [docs] Update old reference to steane_lut decoder by @bmhowe23 in #39
- Add an example of loading a decoder .so by @melody-ren in #42
- Update API documentation by @melody-ren in #71
- Add expressive noise model example by @melody-ren in #123
- Use new CUDA-Q noise models in examples by @bmhowe23 in #120
- [docs] Decoder API clarification by @bmhowe23 in #126
Other Changes
- Update core header files to support C++17 as well by @bmhowe23 in #27
- Add CUDA build support to libraries by @bmhowe23 in #32
- Remove CUDA-Q patch that is no longer needed by @bmhowe23 in #48
- Update CI for arm64 build by @bmhowe23 in #54
- Update test CMakeLists.txt for upstream CUDA-Q change by @bmhowe23 in #93
- Update CUDA-Q build patch by @bmhowe23 in #95
- Test examples in workflow by @melody-ren in #112
What's Changed for Solvers
In addition to arm64 / aarch64 support, this release includes a support for the Bravyi-Kitaev transformation.
Features and Enhancements 🎉
- Add a
get_operator_poolfunction in C++ to mirror the Python by @amccaskey in #13 - Add an option to set the tolerance for jordan_wigner by @melody-ren in #23
- Bravyi-Kitaev implementation by @wsttiger in #35
- Support vector types for kwargs by @bmhowe23 in #99
Bug Fixes 🐛
- Fix signed int overflow in uccsd by @annagrin in #64
- Refactoring and debugging of Jordan Wigner transform (Issue #67) by @kvmto in #82
Breaking Changes 🛠
Documentation Updates ✏️
- Update docs to clarify Python wheels installation requirements by @bmhowe23 in #34
- Update API documentation by @melody-ren in #71
Other Changes
- Update core header files to support C++17 as well by @bmhowe23 in #27
- Add CUDA build support to libraries by @bmhowe23 in #32
- Remove CUDA-Q patch that is no longer needed by @bmhowe23 in #48
- Update CI for arm64 build by @bmhowe23 in #54
- Update test CMakeLists.txt for upstream CUDA-Q change by @bmhowe23 in #93
- Update CUDA-Q build patch by @bmhowe23 in #95
- Test examples in workflow by @melody-ren in #112
New Contributors
- @boschmitt made their first contribution in #8
- @amccaskey made their first contribution in #13
- @melody-ren made their first contribution in #16
- @bmhowe23 made their first contribution in #18
- @wsttiger made their first contribution in #35
- @justinlietz made their first contribution in #7
- @annagrin made their first contribution in #64
- @kvmto made their first contribution in #82
- @caldwellshane made their first contribution in #117
Full Changelog: 0.1.0...0.2.0
0.1.0
This is the initial release of the CUDA-QX libraries! CUDA-QX is a collection of libraries that build upon the CUDA-Q programming model to enable the rapid development of hybrid quantum-classical application code leveraging state-of-the-art CPUs, GPUs, and QPUs. It provides a collection of C++ libraries and Python packages that enable research, development, and application creation for use cases in quantum error correction and hybrid quantum-classical solvers. Please check out the docs and examples for how to get started using the CUDA-QX libraries!
Note: CUDA-QX 0.1.0 depends on CUDA-Q 0.9.0.