0.6.0
CUDA-Q QEC 0.6.0 and CUDA-Q Solvers 0.6.0
This is combined release of CUDA-Q QEC and CUDA-Q Solvers, both version 0.6.0.
This is the first CUDA-Q QEC release that builds example decoder applications on top of CUDA-Q Realtime [blog]. CUDA-Q QEC 0.6 ships with two new real-time-capable decoder pipelines: the RelayBP belief-propagation decoder for qLDPC codes and an NVIDIA Ising convolutional neural network (CNN) pre-decoder paired with a global decoder (PyMatching) for the surface code. These pipelines enable quantum vendors and QEC researchers to deploy real-time GPU decoding for two popular code families via NVQLink.
Additionally, this release of CUDA-Q QEC contains speed improvements for our GPU-accelerated RelayBP decoder (up to 19X!)
For CUDA-Q Solvers 0.6.0, support was added for a new UpCCGSD ansatz solver and a Coupled Exchange Operator (CEO) pool.
Please check out the docs and examples for how to get started using the CUDA-QX libraries!
Note: CUDA-Q QEC 0.6.0 and CUDA-Q Solvers 0.6.0 both depend on CUDA-Q 0.14. For CUDA-Q Realtime usage (experimental), you need to use CUDA-Q 0.14.1.
Features and Enhancements (QEC) 🎉
- Sliding window optimize by @cketcham2333 in #343
- feat(qec): add trt_decoder_config for real-time decoding by @wsttiger in #384
- Add trt cudagraph by @wsttiger in #369
- Add decode batch by @wsttiger in #383
- Create PyMatching decoder plugin by @bmhowe23 in #396
- Add optional
Oparameter to PyMatching plugin by @bmhowe23 in #449 - Update
trt_decoderto supportuint8data types for I/O by @bmhowe23 in #455 - Add graph capture functions to common decoder interface by @bmhowe23 in #475
- Follow-up to #475 - additional decoder interface updates by @bmhowe23 in #478
- Add Hololink QLDPC graph decode bridge and CI test by @cketcham2333 in #481
- Add realtime AI decoder / predecoder infrastructure (GPU + Host) w/ host dispatcher by @wsttiger in #457
- Add FPGA-based test application for realtime predecoder by @wsttiger in #490
nv-qldpc-decoder Updates (Closed Source)
-
Implemented new graph capture interface functions to run RelayBP with CUDA-Q Realtime
-
Add new
repeatableconfiguration option to enable bit-for-bit repeatable results when running back-to-back on the same system -
Significant RelayBP optimizations, for both
fp32andfp64. Timings below show speedups relative to 0.5 for some well-known Bicycle Bivariate Codes on B200. All of the reported speedups are for non-batched, serial execution mode.Case Name (n_k_d) Variant Total Speedup (Ratio) 72_12_6 fp32 3.44 72_12_6 fp64 2.74 144_12_12 fp32 5.50 144_12_12 fp64 4.51 288_12_18 fp32 19.06 288_12_18 fp64 13.16 Average 8.07
Bug Fixes (QEC) 🐛
- Coverity fixes by @cketcham2333 in #395
- Fix OOB r/w issues by @kaiqiy-nv in #464
- Add onnxscript to trt_decoder optional dependency by @wsttiger in #407
- Bug fix and add test cases by @kaiqiy-nv in #424
- Fix pytorch AcceleratorError root-caused by QEC by @kaiqiy-nv in #472
Features and Enhancements (Solvers) 🎉
- UpCCGSD ansatz solver by @rr637 in #372
- Add Coupled Exchange Operator (CEO) pool by @jgonthier in #387
Bug Fixes (Solvers) 🐛
- Fix OOB r/w issues by @kaiqiy-nv in #464
- Fix the gradient evaluation bugs and add test cases by @kaiqiy-nv in #434
- Fix optimiser forwarding bug and other minor bugs and add test cases by @kaiqiy-nv in #435
- Fix BK transformation | Fix JW parity Z-chain | Add test cases by @kaiqiy-nv in #460
- Fix GQE invalid CUDA handle / AcceleratorError when moving model to GPU by @vedika-saravanan in #473
- Fix mixer forwarding | Fix MPI implementation | Add test cases by @kaiqiy-nv in #452
- Add error message for missing system dependencies by @vedika-saravanan in #467
- Fix qubit indices bug and add test case by @kaiqiy-nv in #469
- GQE: PyTorch GPU compatibility check, exit/skip test on mismatch by @vedika-saravanan in #494
Documentation ✏️
- Update documentation for relay-bp by @melody-ren in #346
- [Docs] Add uccgsd to doc by @marwafar in #340
- [Docs] update gen_ham with UHF by @marwafar in #339
- [docs] Update nv-qldpc-decoder docs to describe the new proc_float option by @bmhowe23 in #288
- Incorporate Sliding Window Decoder docs by @bmhowe23 in #359
- Add docs for realtime decoding by @kvmto in #345
- Add docs for AI decoder training with PyTorch by @wsttiger in #344
- fix typo in calling operator pool with uccsd in doc by @marwafar in #366
- Add requirement for memory BP methods in docs by @melody-ren in #376
- Added trt_decoder docs for Python and C++ by @wsttiger in #381
Common / Misc
- License updates by @bmhowe23 in #360
- Add license agreement notification to Docker image by @bmhowe23 in #362
- [core] Fix pre-existing extension point issue by @bmhowe23 in #374
- Bump CUDA-Q commit (with support for breaking changes) by @github-actions[bot] in #416
- Bump CUDA-Q commit (with non-trivial updates) by @github-actions[bot] in #437
- Bump CUDA-Q commit and re-enable some tests by @bmhowe23 in #450
- Bump CUDA-Q dependencies from 0.13 to 0.14 by @bmhowe23 in #468
- Align CUDA-Q and CUDA-Q Realtime commits for 0.14.1 by @bmhowe23 in #489
- Fix minor issues reported by Coverity by @kaiqiy-nv in #399
- Redundantly including Logger.h and FmtCore.h includes ahead of runtime refactor. by @Renaud-K in #409
- Update cuda-quantum-devdeps:ext-... to cuda-quantum-devcontainer-... by @bmhowe23 in #420
- Fix build if CUDAQ_REALTIME_ROOT is not set by @bmhowe23 in #432
- Follow-up to #396 and #416 - fix wheel builds by @bmhowe23 in #439
- Update heterogeneous_map to recognize ints as bools by @bmhowe23 in #441
- Update CMake for TensorRT decoder unit test by @bmhowe23 in #448
Testing
- Update test scripts to allow easy data generation by @bmhowe23 in #354
- Follow-up to #344 - add onnxscript to wheels test env by @bmhowe23 in #368
- [uccsd/uccgsd] Update test tolerances by @bmhowe23 in #375
- Add playback/record to the surface code 1 test by @cketcham2333 in #406
- Add nv-qldpc-test option to surface_code-1.cpp by @cketcham2333 in #415
- Mock decoder by @cketcham2333 in #423
- Add cuda graph launch to mock decoder and introduce autonomous decoder CRTP by @cketcham2333 in #429
- [ci] Update how to get cudaq::realtime by @bmhowe23 in #443
- [realtime] Mock decoder updates for latest cudaq::realtime updates by @bmhowe23 in #444
- [realtime] Move mock decoders to test directories by @bmhowe23 in #445
- [realtime] Advance cudaq::realtime commit and update names accordingly by @bmhowe23 in #446
- Skip GQE GPU tests on unsupported GPU architectures by @vedika-saravanan in #477
- Update container validation script by @bmhowe23 in #487
- Remove cudaq-realtime from All libs CI and enable QLDPC graph test in Release CI by @cketcham2333 in #482
New Contributors
- @kaiqiy-nv made their first contribution in #399
- @rr637 made their first contribution in #372
- @jgonthier made their first contribution in #387
- @vedika-saravanan made their first contribution in #467
Full Changelog: 0.5.0...0.6.0