Repository navigation
CCCL 3.5 Release
The CCCL team is excited to announce the 3.5 release of the CUDA Core Compute Libraries (CCCL). Highlights include public tuning APIs for CUB, reproducible floating-point scans, batched Top-K selection, and reductions that accept problem sizes stored on the GPU.
Public CUB Tuning APIs
CCCL 3.5 exposes public tuning policies for CUB device-wide algorithms. Applications can customize parameters such as threads per block, items per thread, vectorization, and algorithm selection by passing a policy selector through cuda::execution::tune(...) in an execution environment.
Policy selectors choose tuning parameters for a GPU compute capability and return an algorithm-specific policy such as cub::ReducePolicy or cub::ScanPolicy. This provides a supported way to specialize CUB for a workload through the public device-wide APIs.
For example, this example policy below uses 256 threads per block and selects the number of items per thread based on compute capability:
struct ReduceTuning {
__host__ __device__ constexpr cub::ReducePolicy
operator()(cuda::compute_capability cc) const {
auto pass = cub::ReducePassPolicy{
.threads_per_block = 256,
.items_per_thread = cc >= cuda::compute_capability{10, 0} ? 20 : 16,
.vec_size = 4,
.reduce_algorithm = cub::BLOCK_REDUCE_WARP_REDUCTIONS,
.load_modifier = cub::LOAD_LDG
};
return {.multi_tile = pass, .single_tile = pass};
}
};
auto env = cuda::std::execution::env{
cuda::stream_ref{stream}, cuda::execution::tune(ReduceTuning{})
};
auto status = cub::DeviceReduce::Sum(input.begin(), output.begin(), input.size(), env);See the CUB tuning documentation for an example and the CUB environment documentation for more information on how to use environments. Each CUB device-wide algorithm also has a documentation example for how to customize its tuning, see for example how to tune cub::DeviceReduce.
Reproducible Floating Point Scans
cub::DeviceScan now supports run-to-run reproducibility for floating-point summation. Requesting cuda::execution::determinism::run_to_run selects a fixed reduction order, producing repeatable results on the same GPU with the same input and build configuration.
This applies to inclusive and exclusive sums and scans using cuda::std::plus.
See the cub::DeviceScan documentation.
auto env = cuda::std::execution::env{
cuda::stream_ref{stream},
cuda::execution::require(cuda::execution::determinism::run_to_run)
};
auto status = cub::DeviceScan::InclusiveSum(input.begin(), output.begin(), input.size(), env);Batched Top-K Selection
CCCL 3.5 adds cub::DeviceBatchedTopK, which selects the smallest or largest K items independently from many segments. MinKeys, MaxKeys, MinPairs, and MaxPairs support fixed or variable segment sizes and a K value that can vary by segment.
In CCCL 3.5, each segment is processed in one thread block, with a 48 KiB shared-memory limit. With the default policies, the maximum is 8,192 float keys or 4,096 double keys per segment; limits for key-value pairs depend on both types.
Callers must provide a compile-time upper bound on segment size (for example, cuda::args::bounds<1, 8192>() for float keys) and explicitly request non-deterministic, unordered output using determinism::not_guaranteed, tie_break::unspecified, and output_ordering::unsorted through cuda::execution::require(...).
Select the two largest keys per four-element segment from existing device buffers:
constexpr int segment_size = 4, k = 2;
auto segments_in = cuda::make_strided_iterator(
cuda::make_counting_iterator(input.begin()), segment_size);
auto segments_out = cuda::make_strided_iterator(
cuda::make_counting_iterator(output.begin()), k);
auto env = cuda::std::execution::env{
cuda::stream_ref{stream},
cuda::execution::require(
cuda::execution::determinism::not_guaranteed,
cuda::execution::tie_break::unspecified,
cuda::execution::output_ordering::unsorted)
};
auto status = cub::DeviceBatchedTopK::MaxKeys(
segments_in, segments_out,
cuda::args::constant<segment_size>{}, cuda::args::constant<k>{}, num_segments, env);See the CCCL 3.5 cub::DeviceBatchedTopK API for argument annotations, supported types, and complete examples.
CUB Arguments Stored on the GPU
The new <cuda/argument> header provides cuda::args::constant, immediate, deferred, and deferred_sequence, together with argument bounds. These annotations describe values known at compile time, values supplied by the host, and values read from device-accessible memory in stream order.
cub::DeviceReduce::{Reduce, Sum, Min, Max, TransformReduce} now accept a problem size supplied through cuda::args::deferred. A preceding kernel can produce the element count without a round trip to the host. Captured reductions can consume a different count on each CUDA Graph replay without updating or recapturing the graph.
Selection writes its output count to a one-element device buffer. The reduction consumes that count directly on the same stream:
auto selected = cuda::make_device_buffer<float>(stream, device, input.size(), cuda::no_init);
auto count = cuda::make_device_buffer<int>(stream, device, 1, cuda::no_init);
auto sum = cuda::make_device_buffer<float>(stream, device, 1, cuda::no_init);
auto env = cuda::std::execution::env{cuda::stream_ref{stream}};
auto status = cub::DeviceSelect::If(
input.begin(), selected.begin(), count.begin(), input.size(), predicate, env);
status = cub::DeviceReduce::Sum(
selected.begin(), sum.begin(), cuda::args::deferred{count.begin()}, env);cub::DeviceScan::InclusiveScan also accepts an initial value supplied through cuda::args::deferred.
See the DeviceReduce documentation, the InclusiveScan addition, and #8789: support for device-resident problem sizes.
Faster Searches for Sorted Query Values
cub::DeviceFind::LowerBoundSortedValues and UpperBoundSortedValues accelerate batched lower- and upper-bound searches when both the searched range and the query values are sorted under the same comparator. They use a merge-path traversal with O(N + M) total work, where N is the range length and M is the number of queries.
See the cub::DeviceFind documentation.
Additional Library Improvements
- Added
cuda::std::ranges::zip_viewandcuda::std::views::zip, with expanded tuple-like construction and assignment support incuda::std::tupleandcuda::std::pair. - Added
bit_compress,bit_expand,bit_repeat, andbit_reverse, plusshlandshr, to<cuda/std/bit>. - Added
cuda::isclosefor approximate comparisons,cuda::ceil_ilog10, andcuda::device::warp_match_anyfor matching trivially copyable values across warp lanes. - Added
atomic_ref::address()to bothcuda::atomic_refandcuda::std::atomic_ref, andcuda::std::is_virtual_base_ofwhere supported by the compiler. - Added initializer-list overloads for buffer factories, additional memory-pool attribute queries with CUDA 13.3, and updated PTX wrappers for CUDA 13.4.
- Expanded CUB's use of hardware warp-reduction instructions, including floating-point min/max on supported Blackwell targets; added Programmatic Dependent Launch support to
DeviceRadixSortand vectorizedBlockLoad/BlockStoreoperations for contiguous iterators. - Added GDB and LLDB pretty-printers for
cuda::buffer,cuda::std::array, andcuda::std::complexandcuda::complex. - C++23 extended
cuda::std::tupleinterface for tuple like types - Various compile time improvements for type traits
- Added
cuda::std::expected::has_error() - Public macros to detect the host architecture
CCCL_HOST_ARCH
Notable Fixes
- Fixed a race condition in
DeviceTopK. - Fixed out-of-bounds writes in
DeviceHistogramfor large inputs and in streaming reduce-by-key and run-length encoding for single-partition inputs. - Fixed Thrust OpenMP scans processing unused block sums, and improved
cuda::device_bufferiterator compatibility in Thrust andDeviceAdjacentDifference. - Fixed
cuda::std::aligned_allocargument handling,cuda::bufferinitialization for three-byte element types, andto_charswidth calculation for exact powers of ten.
Deprecations and Migration Notes
cub::ChainedPolicy, the CUB dispatch structs and the agent policy types listed below are deprecated. Custom tuning should use policy selectors through the execution environment of the corresponding public Device* API. Both single-call environment APIs and traditional two-phase temporary-storage APIs remain supported. See the tuning migration examples.
All type names in the table are in namespace cub.
| Deprecated customization types | Public tuning policy |
|---|---|
DispatchAdjacentDifference, AgentAdjacentDifferencePolicy |
AdjacentDifferencePolicy |
AgentBatchMemcpyPolicy |
BatchedCopyPolicy |
DispatchHistogram, AgentHistogramPolicy |
HistogramPolicy |
DispatchMergeSort, AgentMergeSortPolicy |
MergeSortPolicy |
DispatchRadixSort, AgentRadixSortDownsweepPolicy, AgentRadixSortUpsweepPolicy, AgentRadixSortHistogramPolicy, AgentRadixSortExclusiveSumPolicy, AgentRadixSortOnesweepPolicy |
RadixSortPolicy |
DispatchReduce, DispatchTransformReduce, AgentReducePolicy |
ReducePolicy |
DispatchReduceByKey, AgentReduceByKeyPolicy |
ReduceByKeyPolicy |
DeviceRleDispatch, AgentRlePolicy |
RleEncodePolicy or RleNonTrivialRunsPolicy, depending on the operation |
DispatchScan, AgentScanPolicy |
ScanPolicy |
DispatchScanByKey, AgentScanByKeyPolicy |
ScanByKeyPolicy |
DispatchSegmentedRadixSort |
SegmentedRadixSortPolicy |
DispatchSegmentedReduce, AgentWarpReducePolicy |
SegmentedReducePolicy |
DispatchSegmentedSort, AgentSubWarpMergeSortPolicy |
SegmentedSortPolicy |
DispatchSelectIf, AgentSelectIfPolicy |
SelectPolicy or PartitionPolicy, depending on the operation |
DispatchThreeWayPartitionIf, AgentThreeWayPartitionPolicy |
ThreeWayPartitionPolicy |
DispatchUniqueByKey, AgentUniqueByKeyPolicy |
UniqueByKeyPolicy |
cub::DeviceCount,DeviceCountUncached, andDeviceCountCachedValueare deprecated. Usecuda::devices.size().cuda::std::inplace_vector::try_push_backandtry_emplace_backnow returncuda::std::optional<T&>instead ofT*. Update code that stores the result in a pointer or compares it withnullptr.- CUB's umbrella and unsupported device-wide headers now issue explicit diagnostics under NVRTC. Include the specific block-, warp-, or thread-level headers used by runtime-compiled kernels.
Full changelog from v3.4.0 to the 3.5 release branch
What's Changed
🚀 Thrust / CUB
📚 Libcudacxx
- [libcu++][doc] Fix Broken
cuda/functionaldocumentation by @fbusato in #9073 cuda::std::simdComplex by @fbusato in #8475libcudacxx-testSKILL by @fbusato in #9100- Fix
mdspanlayout_strideABI failure in MSVC by @fbusato in #8954 cuda::std::simdload and store functionalities by @fbusato in #8252- Update
libcudacxx-styleSKILL by @fbusato in #9115 std::simdpermute by @fbusato in #8508cuda::std::simdBit by @fbusato in #8704cuda::std::simdCreation by @fbusato in #8653cuda::std::simdAlways use_CCCL_HOST_DEVICE_APIby @fbusato in #9191- Fix/Improve
warp_match_allby @fbusato in #9192 cuda::std::simdAlgorithms by @fbusato in #8659- CCCL works with old DLPack versions by @fbusato in #9346
- Optimize
cuda::std::rotl/rotrby @fbusato in #9352 cuda::device::warp_match_anyby @fbusato in #9243- Introduce public
CCCL_HOST_ARCHby @fbusato in #9494 - Support
__float128forisfinite()andisinf()by @fbusato in #9576 - Don't use
__in,__out,__inoutvariable names by @fbusato in #9599 - Use precise header in
bit_castby @fbusato in #9600 cuda::std::simdMemory permute by @fbusato in #8539- Add
cuda::ceil_ilog10by @fbusato in #9613 cuda::std::simdMath by @fbusato in #8740- fix
shared_memory_mdspandocumentation version by @fbusato in #9734 - Workaround for
cuda::std::simdNVRTC c++17 failure with some math functions by @fbusato in #9754 - Update
__cccl_ptx_isafor CTK 13.4 by @fbusato in #9791 cuda::std::simdOptimize small integer operations by @fbusato in #8873cuda::isclose()by @fbusato in #9577cuda::std::simdF32x2 cleanup/refactoring by @fbusato in #8951- Fix
cuda::ptxshrandbfindlong,long longhandling by @fbusato in #10000 cuda::std::simd: Addshr,shl,bit_reverseby @fbusato in #10058cuda::std::simdOptimize Min/Max by @fbusato in #8949
🔄 Other Changes
- Use signed offset type for DevicePartition by @bernhardmgruber in #8971
- cudax: introduce HLL Policy template parameter by @sleeepyjack in #8857
- [libcu++] Turn memory resource properties docs into a table by @pciolkosz in #8973
- Skip header testing for the NVTX3 header. by @wmaxey in #8986
- docs: fix duplicated words in host_stub_visibility and permutation_iterator comments by @vip892766gma in #8980
- Enable clang-diagnositc clang-tidy checks by @Jacobfaib in #8939
- [libcu++] Fix
__detectably_invalidvalue returned during constant evaluation by @davebayer in #8987 - Use the new tuning API for
detail::radix_sort::dispatchby @bernhardmgruber in #7949 - [c2h] Make catch macros work on device by @davebayer in #8928
- Fix nightly GCC8 build regressions in bit_cast and DeviceReduce by @alliepiper in #8924
- [infra] Add clang-cuda-21 in C++23 job for libcu++ by @davebayer in #8988
- [Tile] avoid returns in switch statements in logarithm by @miscco in #9001
- [Tile] Avoid returns inside loops in
mdspanby @miscco in #9004 - [Tile] avoid return in switch statement in
assume_alignedby @miscco in #9002 - [Tile] Avoid returning in a loop in
envby @miscco in #9007 - [Tile] Avoid return in loop in
charconvby @miscco in #9010 - [Tile] avoid use of CPO in variant by @miscco in #9012
- [Tile] Mark
simdas_CCCL_HOST_DEVICEby @miscco in #9013 - [Tile] Avoid return in loop in
variantby @miscco in #9011 - [Tile] Mark threading support as
_CCCL_HOST_DEVICE_APIby @miscco in #9006 - [Tile] Avoid unused variable warning by @miscco in #9009
- [Tile] Avoid returns in loops in various algorithms by @miscco in #9014
- [libcu++] Fix device fp128 functions test by @davebayer in #9015
- Revert "[libcu++] Fix concurent return value writes during device testing (#8957)" by @miscco in #9003
- [Tile] Make
arch_id_CCCL_HOST_DEVICE_APIby @miscco in #9005 - [Tile] Avoid returns in loops in string functions by @miscco in #9008
- [libcu++] avoid warning about pointless comparison by @miscco in #9016
- [Tile] Avoid unused variable warnings by @miscco in #9017
- [Tile] Mark some algorithms as
_CCCL_HOST_DEVICE_APIby @miscco in #9018 - Fix mdspan support in modernized CUB interfaces by @pauleonix in #9020
- [Tile] Avoid use of
in_placeglobal by @miscco in #9024 - [Tile] Avoid returns in loops in some algorithms by @miscco in #9025
- [Tile] Avoid return statements in loops in
bitsetby @miscco in #9027 - [Tile] Do not access global in tile mode by @miscco in #9026
- [Tile] Avoid internal usage of CPOs by @miscco in #9023
- Use the new tuning API internally for
detail::segmented_sort::dispatchby @bernhardmgruber in #8992 - [libcu++] Cleanup our lit config by @miscco in #8917
- [Tile] Avoid calling a device only function in tile mode by @miscco in #9031
- [Tile] Mark format tests as unsupported by @miscco in #9033
- [Tile] Avoid return in switch statements in
variantby @miscco in #9032 - Use also a SM120 GPU in light CI for CCCL.C by @bernhardmgruber in #8972
- Drop override accum_t from reduce by key by @bernhardmgruber in #8993
- [Tile] Try to appease tile compiler by @miscco in #9035
- [Tile] enable some more tests by @miscco in #9034
- [libcu++] Fix default make_shared_resource construction by @bdice in #9044
- [infra] Update nvhpc to 26.3 by @davebayer in #9048
- Fix segmented radix sort benchmark segment size type by @bernhardmgruber in #9039
- Use the new tuning API internally for
detail::segmented_radix_sort::dispatchby @bernhardmgruber in #8927 - Use the new tuning API internally for
detail::select::dispatchandDeviceSelectby @bernhardmgruber in #8880 - Use
U32inby_keybenchmark by @bernhardmgruber in #9051 - Drop obsolete function by @bernhardmgruber in #9050
- Use the new tuning API internally for
detail::select|three_way_partition::dispatchandDevicePartitionby @bernhardmgruber in #8925 - Add work planning how-to by @jrhemstad in #9042
- Use the new tuning API internally for
detail::reduce_by_key::dispatchby @bernhardmgruber in #8756 - Improve argument names in cub/util_math by @pauleonix in #9053
- Complex tanh (and thus tan) accuracy refinement by @s-oboyle in #9036
- Vectorize contiguous iterators in
cub::BlockLoad/Storeby @bernhardmgruber in #9056 - [STF] Add per-handle exec_place stream resources by @caugonnet in #8905
- [cub] Replace
__NVCOMPILER_CUDA_ARCH__withNV_TARGET_MINIMUM_SM_INTEGERby @davebayer in #9067 - Add
__lazy_call_orby @Jacobfaib in #8767 - [libcu++] Always suppress C++ extensions warnings in prologue by @davebayer in #9019
- [libcu++] Explicitly mark
__halfand__nv_bfloat16as trivially_copyable by @miscco in #9069 - [STF] Move unstable_unique from STF to generic cudax utility by @caugonnet in #8190
- [cudax] Disable cuco hashers
__int128test for nvcc 12.0 + gcc combination by @davebayer in #9076 - [cudax] Remove
CUDAX_MEOWtest macros by @davebayer in #8998 - Env passthrough 4/4 by @gonidelis in #9065
- [cub] Replace
assertwithCHECKorREQUIREin tests by @davebayer in #8999 - Add env DeviceTopK without temp_storage args by @gonidelis in #8982
- Final env-passthrough 2/4 by @gonidelis in #8979
- Use the new tuning API internally for
detail::scan_by_key::dispatchby @bernhardmgruber in #8761 - [STF] Make graph_task and stream_task move-only by @caugonnet in #8915
- [STF] Fix graph capture integration by @caugonnet in #8914
- Use public API for deterministic reduction test by @bernhardmgruber in #9087
- Align offset type in unique_by_key to public API by @bernhardmgruber in #9091
- [Tile] improve documentation by @miscco in #9054
- Small segmented radix sort cleanup by @bernhardmgruber in #9092
- Use public API in unique_by_key benchmark by @bernhardmgruber in #9089
- New complex acos function. by @s-oboyle in #9096
- Use the new tuning API internally for
detail::batch_memcpy::dispatchby @bernhardmgruber in #9093 - Final env-passthrough 3/4 by @gonidelis in #8981
- STF: extend cuda_try output-parameter inference (first/last + ambiguity rejection) by @andralex in #8891
- [Tile] Partially support
__halfand__nv_bfloat16in tile mode by @miscco in #9084 - [Tile] Improve table formatting and merge tables by @miscco in #9104
- [Tile] Improve support for
complex,__halfand__nv_bfloat16by @miscco in #9072 - Use the new tuning API internally for
detail::histogram::dispatchby @bernhardmgruber in #9106 - [libcu++] Do not require
default_initializablefor bit_cast by @miscco in #9105 - [CUB] Fix invalid condition to exclude 128 bit integers by @miscco in #9112
- Fix RAPIDS in CI by @trxcllnt in #9103
- Update
.gitignoreto excludecodegraphfiles by @fbusato in #9114 - Remove dead code by @bernhardmgruber in #9122
- [libcu++] Fix redefinition errors in
cuda::std::simdtest utilities by @shwina in #9125 - Distribute
std::tuple_(size|element)specializations to implementation files by @davebayer in #9123 - use
std::scoped_lockby @charan-003 in #9124 - [libcu++] Remove empty messages in
static_assertby @davebayer in #9131 - [STF] Properly destroy CUDA streams and do not try to initialize CUDA while capturing by @caugonnet in #8919
- Refactor DeviceReduce dispatch logic by @bernhardmgruber in #9088
- [STF] [TRIVIAL] Rename decorated_stream to augmented_stream by @andralex in #9138
- Make more warpspeed scan stage counts tunable by @bernhardmgruber in #9128
- Skip plotting benchmarks without data by @bernhardmgruber in #9085
- Use the new tuning API internally for
detail::batched_topk::dispatchby @bernhardmgruber in #9095 - [STF] Use out parameter for partition mappers by @caugonnet in #9117
- Hardcode
DeviceTransformOffsetTtoint64by @bernhardmgruber in #9127 - [STF] Deduplicate cudaStreamIsCapturing helper by @andralex in #9136
- [libcu++] Implement
std::constant_wrapperby @davebayer in #9046 - [libcu++] Change
inplace_vectorreturn type fortry_meowmethods by @davebayer in #9130 - [libcu++] Fix
ranges::subrangeuse ofpair-likeby @miscco in #9126 - [libcu++] Update
__cccl_ptx_isafor CTK 13.3 by @davebayer in #9164 - Support launching the devcontainer from a git worktree by @alliepiper in #8950
- CUDA runtime unit test by @charan-003 in #8859
- [STF] Add C API context wait helper by @caugonnet in #9161
- Use the public API for the
transform_reducebenchmark by @bernhardmgruber in #9179 - [cudax] Define
_CUDAX_ENABLE_GROUP_FEATURES_IN_LIBCUDACXXglobally in cudax by @davebayer in #9180 - Refactor warpspeed scan 1/2 by @bernhardmgruber in #9168
- [cub] Fix warp reduce of fixed size random access range by @davebayer in #9153
- Only enabled PDL for PTX/SASS supporting it by @bernhardmgruber in #9163
- [cudax] Disable cuco hashers test for nvcc 12.0 by @davebayer in #9185
- Update all the things to CTK 13.3. by @alliepiper in #9143
- [STF] Fix host-thread data races in concurrent task submission by @caugonnet in #9186
- [cudax] Make
lane_maskpart of the mapping result by @davebayer in #9140 - [cudax] Initial
cudax::coop::reduceprototype by @davebayer in #9154 - [libcu++] Fix
cuda::std::constant_wrappertests for nvcc 13.3 by @davebayer in #9198 - Refactor warpspeed scan 2/2 by @bernhardmgruber in #9169
- Use public CUB API to implement thrust::unique_by_key by @bernhardmgruber in #9197
- Fix MSVC returning incorrect values for complex tan/tanh for large inputs. by @s-oboyle in #9189
- [cudax] Implement
binary_partitionmapping for threads within warp by @davebayer in #8894 - [cudax] Remove synchronizers header reintroduced during rebasing by @davebayer in #9204
- CCCL C v2 by @shwina in #8985
- Avoid passing OOB data to scan operator in warpspeed scan by @bernhardmgruber in #9202
- smoke test to verify GPU memory allocation/deallocation by @charan-003 in #9195
- [STF] Add extended C context creation API by @caugonnet in #9162
- [cudax] Implement
cudax::coop::reduceforcudax::this_clusterby @davebayer in #9167 - [libcu++] Implement tuple-like constructors for
tupleby @miscco in #9059 - [libcu++] Replace
cudaStreamPerThreadwithcudaStream{}in PSTL by @davebayer in #9214 - [libcu++] Suppress
-Wattributesin lit tests with nvcc 12.0 and gcc by @davebayer in #9216 - Add coderabbit disclaimer by @shwina in #9221
- Argument annotation framework for segmented algorithms by @pciolkosz in #8875
- [STF] Add re-launchable popped graphs to stackable_ctx by @caugonnet in #9178
- [libcu++] Fix use of exception keywords by @davebayer in #9220
- Marks conditionally-used params in argument annotation framework
[[maybe_unused]]by @elstehle in #9225 - Use the new tuning API internally for
detail::segmented_scan::dispatchby @gonidelis in #9194 - Allow public tuning of
cub::DeviceAdjacentDifferenceby @bernhardmgruber in #9218 - Allow public tuning of
cub::DeviceMergeSortby @bernhardmgruber in #8600 - Fix NVRTC 13.3 warning bug in CCCL and CCCL.C by @NaderAlAwar in #9171
cudax::fill_bytes(mdspan)by @fbusato in #9193- cudax/stf: migrate stackable/ from cuda_safe_call to cuda_try by @andralex in #9165
- run_to_run deterministic device scan by @srinivasyadav18 in #9098
- [cub] Replace cub parameter framework with cuda::argument by @pciolkosz in #9074
- [libcu++] Rename max to highest in arguments framework by @pciolkosz in #9246
- [cudax] Implement
cudax::coop::reduceforcudax::this_gridby @davebayer in #9203 - [libcu++] Use stream's context in PSTL by @davebayer in #9219
- Prevent type-mismatch errors in cuda.compute.reduce_into by @acosmicflamingo in #9206
cudax::copy(mdspan)Optimize shared memory cases by @fbusato in #9137- [libcu++] Fix issues with new tuple constructors by @miscco in #9261
- Use the new tuning API internally for detail::find::dispatch by @gonidelis in #9240
- [cudax] Update lane mask inside mappings only when unit is thread by @davebayer in #9264
- Use uniform type names for init values throughout CUB by @gonidelis in #9267
- Build more RAPIDS libraries in CI by @trxcllnt in #9116
- fix: Replace runtime_static_assert tests with compile tests. by @arnavnagzirkar in #9211
- Rename
block_scan_algorithmtoscan_algorithmby @bernhardmgruber in #9244 - [libcu++] Move argument bounds helpers to bound file by @miscco in #9276
- [CUB] Adds benchmarks for batched indexed top-k (aka batched arg top-k) by @elstehle in #9288
- [cudax] Implement
cudax::invoke_oneby @davebayer in #9230 - [thrust] Fix missing qualifiers for basic_common_reference by @miscco in #9292
- [libcu++] Disable SIMD tests for tile mode by @miscco in #9296
- Deprecate
AgentMergeSortPolicyandAgentAdjacentDifferencePolicyby @bernhardmgruber in #9235 - [CUB] Makes the batched top-k selection direction compile-time only by @elstehle in #9286
- Allow public tuning of
cub::DeviceTransformby @bernhardmgruber in #8745 - [cuda.compute] Enable building cuda-compute against CCCL C V2 by @shwina in #9200
- [STF] Add C bindings for the places layer by @caugonnet in #9232
- cudax/stf: migrate internal/ launch + host_launch_scope from cuda_safe_call to cuda_try by @andralex in #9249
- Restore override key with a comment by @shwina in #9298
- Allow public tuning of
cub::DeviceForby @bernhardmgruber in #9297 - Allow public tuning of
DeviceFindby @bernhardmgruber in #9236 - cudax/stf: migrate internal/ misc files from cuda_safe_call to cuda_try by @andralex in #9241
- cudax/stf: migrate internal/ parallel_for + cuda_kernel scopes from cuda_safe_call to cuda_try by @andralex in #9265
- [cudax] Implement
cudax::coop::reducefor warp groups within a block by @davebayer in #9258 - Allow public tuning of non-ByKey
cub::DeviceScanby @bernhardmgruber in #8853 - Support
cub::DeviceReducewithout initial value by @bernhardmgruber in #9289 - Better asserts for warp primitives without temp storage by @bernhardmgruber in #9294
- Fix the references to permutation_iterator in shuffle_iterator docs by @djns99 in #9307
- [libcu++] Fix
__is_sequencedefinition for arguments framework by @miscco in #9271 - [infra] Add
clang-22CUDA build job to nightly/weekly CI by @davebayer in #8888 - [libcu++] Implement tuple protocol for
integer_sequenceby @davebayer in #9129 - Allow public tuning of
cub::DeviceScan::*ByKeyby @bernhardmgruber in #9215 - [STF] Add C bindings for stackable contexts by @caugonnet in #9233
- Also test Thrust on SM120 in nightly/weekly CI by @bernhardmgruber in #9201
- [cudax] Implement
cuda::coop::reducefor threads within a warp by @davebayer in #9300 - Take environments by
const&inDeviceTransformby @bernhardmgruber in #9336 - Document CUB temp storage alignment by @jrhemstad in #9302
- Deprecate
DispatchScanandAgentScanPolicyby @bernhardmgruber in #9234 - Allow public tuning of the
cub::DeviceReduce::ReduceByKeyandcub::DeviceRunLengthEncode::Encodeby @bernhardmgruber in #9329 - [STF] Migrate __stf/allocators/ from cuda_safe_call to cuda_try by @andralex in #9147
- [Tile] Improve testing of
__tile__and__device__only functions by @miscco in #9313 - Refresh c2h inspect_changes fixture path by @caugonnet in #9356
- [libcu++] Implement C++26 new tuple-assignments by @miscco in #9227
- [libcu++] Adds
stable_sortedoutput_ordering by @elstehle in #9355 - Make thrust distributions compatible with URNG interface by @RAMitchell in #9319
- Update documentation for cuda-cccl to say python 3.10+ and CC 7.5+ by @NaderAlAwar in #9365
- Use sccache-dist build cluster in optional third-party jobs by @trxcllnt in #9118
- Rename non-library uses of "warpspeed" to "lookahead" by @bernhardmgruber in #9327
- cudax/stf: migrate internal/ context + resources from cuda_safe_call to cuda_try by @andralex in #9248
- Update devcontainers missed in #9118 by @trxcllnt in #9368
- docs: add missing doxygen comment for par_nosync_t::on() by @Oxygen56 in #9196
- [libcu++] Make argument namespace and wrappers construction public by @pciolkosz in #9251
- [cudax] Implement
cudax::coop::shufflefor threads within a warp by @davebayer in #9325 - [cudax] Implement
cudax::coop::shuffle_downfor threads within a warp by @davebayer in #9371 - [cudax] Use
coop::shuffle_downincoop::reduceby @davebayer in #9392 - [cudax] Implement
cudax::coop::shuffle_upfor threads within a warp by @davebayer in #9390 - Implement
cuda::std::basic_format_stringby @davebayer in #5569 - Override GPU name by @gevtushenko in #9385
- Adds tests for segment-specific-k-values by @elstehle in #9311
- [cudax][STF] Clarify stream pool capture and teardown comments by @caugonnet in #9395
- [libcu++] Fix the default device pool getter by @pciolkosz in #9351
- [cudax][STF] Auto-free graph allocations on launch by @caugonnet in #9394
- Intentionally do not
dlcloseJIT compiled libraries in v2 (hostJIT) by @shwina in #9402 - Improves segmented top-k test compilation times by @elstehle in #9404
- [libcu++] Adds a
cuda::execution::tie_breakrequirement by @elstehle in #9238 - [PSTL] Use env based overload for
DeviceFindIfby @miscco in #9318 - [libcudacxx] Add cuda::std::__stringof for compile-time values including function names by @andralex in #9299
- [libcu++] harden the preprocessor machinery to avoid user defined tokens by @miscco in #9407
- [cuda.compute]: stop wrapping binary search comparator in python callable by @NaderAlAwar in #9428
- Histogram tuning policy cleanup by @bernhardmgruber in #9361
- Allow public tuning of the
cub::DeviceRunLengthEncode::NonTrivialRunsby @bernhardmgruber in #9347 - Allow public tuning of
cub::DeviceSelect(withoutUniqueByKey) andcub::DevicePartition(without three-way) by @bernhardmgruber in #9316 - Remove unused scan output type local by @fallintoplace in #9410
- Fix tuning docs for DeviceFind by @bernhardmgruber in #9379
- Remove __syncthreads() duplicate from agent_radix_histogram by @gonidelis in #9445
- run to run scan warpspeed impl sm100+ by @srinivasyadav18 in #9263
- [libcu++] Fix device memory pool test by @pciolkosz in #9442
- Smoke test to verify pinned memory by @charan-003 in #9285
- Productize tuning API for
DeviceSegmentedScanby @gonidelis in #9430 - Allow public tuning of three-way
cub::DevicePartition::Ifby @bernhardmgruber in #9324 - Refactor libcudacxx-style skill by moving CCCL wide style guidelines to a specific file by @NaderAlAwar in #9405
- [CUB, docs-only] Adds docs page on the requirements users can express for
DeviceTopKandDeviceBatchedTopKby @elstehle in #9446 - [cuda.compute]: cache np.dtypes properly in stateful ops by @NaderAlAwar in #9469
- [cudax] Implement broadcasted variants of
cudax::coop::reduceby @davebayer in #9360 - Update NVBench by @gonidelis in #9223
- Reorganize cuco implementation headers under detail/ by @PointKernel in #9436
- Allow public tuning of
cub::DeviceHistogramby @bernhardmgruber in #9362 - Allow public tuning for
cub::DeviceSelect:::UniqueByKeyby @gonidelis in #9370 - Add PDL to
cub::DeviceRadixSortby @gonidelis in #9247 - Rename
warp_threadsin tuning policies by @bernhardmgruber in #9485 - Enable bugprone clang-tidy checks by @Jacobfaib in #9467
- CRITICAL: Add PDL guard back missed in #9247 by @gonidelis in #9497
- [libcu++] Implement
cuda::std::vformat_toby @davebayer in #9443 - [libcu++] Indirect
indirect_binary invocableby @miscco in #9417 - [libcu++] Implement
cuda::std::formatted_sizeby @davebayer in #9472 - Radix sort policy name improvements by @bernhardmgruber in #9489
- Allow public tuning of
cub::DeviceMemcpyby @bernhardmgruber in #9359 - Ensure warpspeed scan uses <= 256 threads for NVHPC by @bernhardmgruber in #9490
- bugprone-signed-char-misuse by @Jacobfaib in #9507
- bugprone-sizeof-expression by @Jacobfaib in #9517
- bugprone-multi-level-implicit-pointer-conversion by @Jacobfaib in #9509
- bugprone-unhandled-self-assignment by @Jacobfaib in #9522
- bugprone-empty-catch by @Jacobfaib in #9514
- bugprone-assignment-in-if-condition by @Jacobfaib in #9519
- [Tile] Disable tile mode for NVCC 13.3 by @miscco in #9488
- [Tile] Mark alignment helpers as
_CCCL_HOST_DEVICE_APIby @miscco in #9487 - [CUB] Refactor
DeviceSelect::Flaggedto always take an environment by @miscco in #9455 - Wraps min/max in parantheses to avoid MSVC compilation issues by @elstehle in #9537
- bugprone-suspicious-stringview-data-usage by @Jacobfaib in #9525
- bugprone-inc-dec-in-conditions by @Jacobfaib in #9527
- Fix
BlockTopK+/-0.0 handling by @pauleonix in #9470 - [STF] Make exec_place/data_place singletons thread-safe by @caugonnet in #9541
- [STF] Fix data race on stackable_logical_data across host threads by @caugonnet in #9540
- bugprone-suspicious-include by @Jacobfaib in #9508
- bugprone-forward-declaration-namespace by @Jacobfaib in #9504
- [cccl.c] Split build step into compile + load by @shwina in #8484
- bugprone-unintended-char-ostream-output by @Jacobfaib in #9521
- Add basic communicator concept by @Jacobfaib in #9426
- bugprone-return-const-ref-from-parameter by @Jacobfaib in #9515
- use
__ballot_syncin warpspeed lookahead by @srinivasyadav18 in #9471 - Add a workaround for nvcc bug in constant wrapper and revert CUB tests changes by @pciolkosz in #9382
- Temporarily disable is_device_accessible peer tests by @pciolkosz in #9547
- [HostJit] Properly use
__declspecon windows by @miscco in #9539 - [libcu++] Implement
cuda::std::format_toby @davebayer in #9474 - [libcu++] Skip
__fp_set_expfpclassifytests on denormals by @davebayer in #9536 - [libcu++] Implement
cuda::std::format_to_nby @davebayer in #9482 - Implement
ranges::zip_viewby @Jacobfaib in #8744 - [cuda.compute]: add benchmarks to measure host side overhead by @NaderAlAwar in #9432
- bugprone-undefined-memory-manipulation by @Jacobfaib in #9518
- bugprone-move-forwarding-reference by @Jacobfaib in #9526
- [libcu++] Implement
cuda::std::dynamic_formatby @davebayer in #9483 - Add multi GPU CI job for libcu++ by @pciolkosz in #9435
- bugprone-casting-through-void by @Jacobfaib in #9523
- [libcu++] Guard cuda::args bounds against types without numeric_limits by @edenfunf in #9473
- [libcu++] Consistently waive memory pool tests if unsupported by @pciolkosz in #9479
- [cuda.compute]: add CI job for
minimalcuda-cccl extra by @NaderAlAwar in #9434 - [CUB] Refactor
DeviceAdjacentDifference::SubtractLeftto always take an environment by @miscco in #9418 - [libcu++] Implement
cuda::std::range_formatby @davebayer in #9558 - [libcu++] Implement tuple-like constructors for
pairby @miscco in #9543 - [CUB] Refactor
DeviceAdjacentDifference::SubtractRightto always take an environment by @miscco in #9419 - [STF] Support ctx.wait() on a token by @caugonnet in #9501
- [libcu++] Implement
cuda::std::formattableconcept by @davebayer in #9544 - Add determinism docs by @srinivasyadav18 in #9350
- Fix clang-tidy early return by @Jacobfaib in #9580
- bugprone-unchecked-optional-access by @Jacobfaib in #9520
- [CUB][Bug] Fix DeviceHistogram out-of-bounds write by @fbusato in #9570
- [DOC] Fix sidebar noise from Breathe overload anchors by @gonidelis in #9585
- [libcu++] Fix peer access case in is_pointer_accessible by @pciolkosz in #9478
- [thrust] Use CCCL Runtime in Thrust set operation tests by @pciolkosz in #9551
- bugprone-pointer-arithmetic-on-polymorphic-object by @Jacobfaib in #9528
- [HostJit] Windows support by @miscco in #9502
- bugprone-forwarding-reference-overload by @Jacobfaib in #9512
- bugprone-narrowing-conversions by @Jacobfaib in #9505
- bugprone-use-after-move by @Jacobfaib in #9516
- Refactor
device_adjacent_difference::dispatchto take a tuning by @miscco in #9454 - [CUB] Add DeviceFind lower/upper bound for sorted values via merge-path by @AneeshGidda in #8780
- [CUB] Refactor
cub::DeviceSelect::Ifto always take an environment by @miscco in #9456 - [CUB] Refactor
DeviceSelect::FlaggedIfto always take an environment by @miscco in #9457 - [CUB] Refactor
DeviceSelect::Uniqueto always take an environment by @miscco in #9458 - bugprone-integer-division by @Jacobfaib in #9524
- [cuda.compute]: Relax cache key in histogram by @NaderAlAwar in #9596
- [libcu++] Add initializer_list overloads for make_*_buffer helpers by @pciolkosz in #9586
- Use cuda::buffer in ccclrt algorithm tests by @pciolkosz in #9591
- [libcu++] Optimize integral formatters by @davebayer in #9606
- [libcu++] Implement pair-like assignments for
pairby @miscco in #9579 - [libcu++] Set current context for buffer driver operations by @pciolkosz in #9615
- Move coderabbit disclaimer to PR template and collapse walkthrough by @NaderAlAwar in #9610
- Allow public tuning of
cub::DeviceRadixSortby @bernhardmgruber in #9491 - [libcu++] Optimize
to_charsintegral width calculation by @davebayer in #9601 - [CUB] Refactor
DeviceHistogram::MultiHistogramEvento always take an environment by @miscco in #9552 - [CUB] Refactor
DeviceHistogram::HistogramEvento always take an environment by @miscco in #9553 - [CUB] Refactor
DeviceHistogram::MultiHistogramRangeto always take an environment by @miscco in #9554 - [CUB] Refactor
DeviceHistogram::HistogramRangeto always take an environment by @miscco in #9555 - [libcu++] Add image processing CCCL Runtime / CUB example by @pciolkosz in #8541
- [libcu++] Add tests for cross device APIs and APIs related to a device with a different device set current by @pciolkosz in #9617
- [cccl.c]: Add
serialize()anddeserialize()functions to enable ahead-of-time compilation workflows by @shwina in #9568 - Enforce MergePolicy tuning values by @bernhardmgruber in #9439
- Duplicate radix sort dispatch to deprecate
DispatchRadixSortby @bernhardmgruber in #9530 - [STF] Exempt STF test CMake files from CMake owners by @caugonnet in #9642
- Add NCCL communicator by @Jacobfaib in #9427
- [docs] Add a note about error handling using exceptions to the docs by @pciolkosz in #9632
- [libcu++] Guard tuple and pair against dangling references by @miscco in #9622
- Pacify clang-tidy for ncclCommInitAll() by @Jacobfaib in #9651
- [cuda.compute]: fix v2 issue with some well known ops not being supported by @NaderAlAwar in #9649
- Exposes
DeviceBatchedTopK::{Min,Max}{Keys,Pairs}for non-deterministic, unordered, and small segments-only by @elstehle in #9331 - Unify atomic and two-phase reduction by @bernhardmgruber in #9349
- [CUB] Refactor
DeviceCopyto always take an environment by @miscco in #9416 - [CUB] Cleanup some dispatch arguments by @miscco in #9603
- [cub] Specialize
std::formatterfor enums by @davebayer in #9641 - Move
bits_per_passlast intopk_policyby @bernhardmgruber in #9640 - [cudax] Add
unit_prefix to unit-related data in mapping result by @davebayer in #9625 - Allow public tuning of
cub::DeviceMergeby @bernhardmgruber in #9637 - Allow public tuning of
cub::DeviceCopyby @bernhardmgruber in #9658 - Move CUB test docs into developer docs by @bernhardmgruber in #9662
- Update Node 20 GitHub Actions by @jrhemstad in #9673
- [cudax][cuco] Use the detail config header and proper visibility macros in cuco headers by @PointKernel in #9675
- [STF] Fix slice copy context activation on multi-GPU by @caugonnet in #9677
- [STF] Fix multi-context parallel for for grid places by @caugonnet in #9604
- Improve tuning introduction docs by @bernhardmgruber in #9661
- Unify reduce and rfa policy by @bernhardmgruber in #9639
- [CUB] Refactor
DeviceReduce::Arg{Min, Max}to always take an environment by @miscco in #9403 - Add
DeviceSegmentedScanto device-wide docs by @bernhardmgruber in #9690 - [TRIVIAL] Pin numba below 0.66 for Python CUDA extras by @caugonnet in #9692
- Fix NCCL tests for multi-GPU by @Jacobfaib in #9693
- fix: cache cuda.compute builds for closures over Python scalars by @nethum529 in #9680
- [libcu++] Fix iterator traversal checks for
minmax_elementby @miscco in #9694 - [libcu++] Try and avoid MSVC circular is_constructible chain by @miscco in #9683
- [HostJit] Add all tested device APIs to the PCH cache by @miscco in #9663
- Productize
DeviceSegmentedSorttuning API by @gonidelis in #9681 - Use
cuda::std::is_sufficiently_alignedin CCCL by @davebayer in #9684 - [cudax] Implement
cudax::coop::any_ofalgorithm for <= warp groups by @davebayer in #9665 - [CUB] Refactor
dispatch_streaming_arg_reduceto take a tuning environment by @miscco in #9660 - Allow public tuning of non-ByKey
cub::DeviceReduceby @bernhardmgruber in #8863 - Allow public tuning of
cub::DeviceFind::*BoundSortedValuesby @bernhardmgruber in #9688 - Rename
FindPolicytoFindIfPolicyby @bernhardmgruber in #9689 - Allow public tuning of
cub::DeviceSegmentedRadixSortby @bernhardmgruber in #9685 - Fix
cudafe++< 13.1 with older gcc by @davebayer in #9657 - Fix CUB DeviceReduce env overloads to accept no_init_t by @Jacobfaib in #9676
- Fix use of
reduce_policyby @davebayer in #9704 - Add fixed_capacity_map to cudax by @srinivasyadav18 in #7705
- [libcu++] Additional peer device copy testing by @pciolkosz in #9636
- [places] Align cyclic_shape::size with iteration cardinality by @fallintoplace in #9148
- Allow public tuning of
cub::DeviceSegmentedReduceby @bernhardmgruber in #9686 - Small refactorings for batched memcpy by @bernhardmgruber in #9659
- Improve NV_IF_ELSE_TARGET formatting in segmented reduce dispatch by @bernhardmgruber in #9709
- Control unrolling in
cub::DeviceMerge[Sort]via tuning by @bernhardmgruber in #9181 - Remove
num_from tuning policy members by @bernhardmgruber in #9717 - [cuda.compute] Expose
.serialize()and.deserialize()methods in Python by @shwina in #9644 - bugprone-misplaced-widening-cast by @Jacobfaib in #9506
- Rename
AgentTopKPolicytoagent_topk_policyby @bernhardmgruber in #9711 - Fix doxygen tuning doc headings by @bernhardmgruber in #9713
- Rename policy selector tests by @bernhardmgruber in #9715
- [STF][trivial] Forward declare stf_ctx_handle in C-STF header place section by @caugonnet in #9705
- [docs] Document that streams are created as non-blocking by @pciolkosz in #9707
- Fix cuco clang-tidy error by @Jacobfaib in #9736
- Vectorize output store in ublkcp DeviceTransform kernel by @nanan-nvidia in #9481
- Separate from and deprecate
DispatchScanByKeyby @bernhardmgruber in #9714 - [cudax] Change cuco test target names by @davebayer in #9738
- [cub] Fix clang-tidy narrowing conversion warning by @davebayer in #9741
- Swap
small_segmentandmedium_segmentinSegmentedSortPolicyby @bernhardmgruber in #9716 - Cover __half and __nv_bfloat16 in CUB Reduce/Scan/RadixSort benchmarks by @edenfunf in #9708
- Make stream and memory resource explicit in fixed_capacity_map APIs by @PointKernel in #9719
- [CUB] warpspeed kernel: make scan_resources_t copy by @srinivasyadav18 in #9751
- [cuda.compute]: Enable (AoT) compilation for multiple compute capabilities by @shwina in #9732
- Rename cuco
.hppheaders to.cuhby @PointKernel in #9735 - [STF] Name green context data places by handle and support VMM mem_create by @caugonnet in #9706
- Fix test helper producing inverted interval by @pauleonix in #9753
- [STF] Export blocked_partition_custom into cuda::experimental::stf by @caugonnet in #9758
- Make stream and memory resource explicit in hyperloglog APIs by @PointKernel in #9720
- warpspeed run_to_run deterministic scan for SM90 using atomic global counter by @srinivasyadav18 in #9565
- [CUB]
DeviceReducewith device resident problem size by @NaderAlAwar in #9722 - Symlink agent files instead of telling bots to read other files by @Jacobfaib in #9750
- Deprecate
cub::ChainedPolicyby @bernhardmgruber in #9744 - Deprecate CUB device count functions by @bernhardmgruber in #9743
- Make DeviceTransform a P0 benchmark for QA by @bernhardmgruber in #9774
- Do not vectorize large type reductions by @bernhardmgruber in #9762
- Tuning policy cleanup 1/2 by @bernhardmgruber in #9710
- Strip zero-width or unprintable unicode characters by @Jacobfaib in #9564
- [Tile] tile DeviceTransform port by @nanan-nvidia in #9210
- Create a swapfile in
workflow-run-job-linuxby @trxcllnt in #9739 - Test unaligned temp storage by @bernhardmgruber in #9746
- Run dummy test if unit test would be empty in C++17 by @bernhardmgruber in #9784
- Fully classify the
cub::detail::InputValuetemplate to avoid implicit conversions by @Jacobfaib in #9772 - [cuda.compute][cuda.coop]: Replace all usages of device arrays outside examples with new wrapper hat does not depend on cupy or numba-cuda by @NaderAlAwar in #9653
- Handle a type of size exactly 3 in cuda::buffer by @Jacobfaib in #9776
- Tuning policy cleanup 2/2 by @bernhardmgruber in #9712
- bugprone-branch-clone by @Jacobfaib in #9513
- Add builds for MSVC cccl_c_parallel by @miscco in #9605
- Use __builtin_bswapg in cuda::std::byteswap when available by @Functionhx in #9785
- [infra] Add multi-gpu CI for cudax by @pciolkosz in #9590
- Add MGMN Reduce by @Jacobfaib in #9645
- [libcu++] Use cccl runtime in thrust tests pt 2 by @pciolkosz in #9633
- Remove
_CCCL_GRID_CONSTANTfromDeviceSelectSweepKernelparameters by @nanan-nvidia in #9795 - Fix Thrust contiguous iterator unwraps for cuda::device_buffer by @sleeepyjack in #9756
- [STF] Implement cyclic_partition::get_executor by @caugonnet in #9803
- Add tuning policies to CUB API docs by @bernhardmgruber in #9745
- Nits in
score.pyby @gonidelis in #9809 - Fix scanning out-of-bounds items in OpenMP scan by @bernhardmgruber in #9759
- [docs] Fix saturating overflow arithmetic docs by @davebayer in #9812
- Reduce P0 DeviceTransform benchmarks to babelstream and fill by @bernhardmgruber in #9780
- Pin NVTX to commit before upstream scope refactor by @bernhardmgruber in #9820
- [cub] Fix
bugprone-branch-cloneerror by @davebayer in https://github.com/NVIDIA/cccl/pull/9810 - [cudax][cuco] Port fixed_capacity_map benchmarks from cuCollections by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/9748
- Remove _CCCL_GRID_CONSTANT from merge sort kernel parameters by @nanan-nvidia in https://github.com/NVIDIA/cccl/pull/9829
- [libcu++] Implement P3798R1 The unexpected in
std::expectedby @davebayer in #9733 - Refactor enum to string functions by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9752
- [cudax] Implement
cudax::takemapping by @davebayer in https://github.com/NVIDIA/cccl/pull/9818 - Use plus tunings for lookahead scan more widely by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9813
- Move cccl_get_catch2() into testing block by @codingwithmagga in https://github.com/NVIDIA/cccl/pull/9383
- [clang-tidy] Suppress new warnings emitted by clang-tidy-22 by @davebayer in https://github.com/NVIDIA/cccl/pull/9855
- [cccl.c] Fix scan policy mismatch between host and NVRTC by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9770
- DeviceScan::InclusiveScan() accepts a
cuda::args::deferredby @Jacobfaib in #9826 - Use
cuda::std::numbersin cccl by @davebayer in https://github.com/NVIDIA/cccl/pull/4955 - Remove most
CCCL_GRID_CONSTANTannotations by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9858 - Revert "Pin NVTX to commit before upstream scope refactor" (#9820) by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9850
- Add multi-GPU exclusive_scan by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9796
- Add deferred argument to 2-phase InclusiveScan by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9865
- [CUB] GPU-to-GPU
DeviceReducewith device resident problem size by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9740 - [FEA] Replace not_equal_to_val with cuda::std::not_fn(cuda::equal_to_value{}) by @rkothari3 in https://github.com/NVIDIA/cccl/pull/9833
- Move SASS diff description to a SKILL by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9767
- [Thrust] Split transform test by @miscco in https://github.com/NVIDIA/cccl/pull/9701
- Add pytorch inspired DeviceTransform benchmark by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9764
- Add
cudax::distributedandcudax::returned_tospecifiers for cooperative algorithms by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9862 - Remove unnecessary
is_evenoverloads for complex by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9873 - Replace
bool_constantbyif constexprin agent_sub_warp_merge_sort by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9876 - [cudax] Fix cudax benchmarks prefix by @davebayer in https://github.com/NVIDIA/cccl/pull/9878
- [STF] Add exec places from externally-owned CUDA contexts by @caugonnet in https://github.com/NVIDIA/cccl/pull/9779
- [CUB] Refactor
DeviceSelect::UniqueByKeyto always take an environment by @miscco in https://github.com/NVIDIA/cccl/pull/9459 - [cuda.compute]: make cuda.compute thread safe to enable free threaded wheels by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9475
- [CUB] Refactor
DevicePartition::Flaggedto always take an environment by @miscco in https://github.com/NVIDIA/cccl/pull/9463 - Add multi-GPU inclusive_scan by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9827
- Add
cstdioandcstdarghostlib headers by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9879 - [cudax][cuco] Add HyperLogLog benchmarks by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/9889
- [cuda.cccl] Simplify Python CMake configuration by @caugonnet in https://github.com/NVIDIA/cccl/pull/9883
- [CUB] Refactor
DevicePartition::Ifto always take an environment by @miscco in https://github.com/NVIDIA/cccl/pull/9464 - [places] Add checked execution grid reshaping by @caugonnet in https://github.com/NVIDIA/cccl/pull/9898
- Fix dead 64-bit rotate builtin macros and popcount tile undef in by @temujinkz in https://github.com/NVIDIA/cccl/pull/9874
- [STF] Add structured tensor partitions and placement evaluation to the places layer by @caugonnet in https://github.com/NVIDIA/cccl/pull/9804
- [libcu++] Add
__float128support forcuda::std::fabsby @davebayer in https://github.com/NVIDIA/cccl/pull/9895 - Don't use the thrust size dispatchers anymore in multi-GPU by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9978
- Fix tuning policy links in docs by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9897
- Slightly bump test sizes for mgmn tests by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9980
- Add initial docs for CI + infrastructure systems by @alliepiper in https://github.com/NVIDIA/cccl/pull/9611
- cudax/stf: migrate stream/interfaces/ from cuda_safe_call to cuda_try by @andralex in https://github.com/NVIDIA/cccl/pull/9268
- [STF] Drive parallel_for and data placement from structured partitions by @caugonnet in https://github.com/NVIDIA/cccl/pull/9808
- Use result policies for MGMN algorithms by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9979
- Add unaligned versions of cub::DeviceTransform babelstream and fill benchmarks by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9881
- Add owning nccl communicator by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9977
- [bench] Enable metatargets for benchmark targets by @davebayer in https://github.com/NVIDIA/cccl/pull/9990
- cudax/stf: cuda_try migration — stream event_types (PR7) by @andralex in https://github.com/NVIDIA/cccl/pull/9303
- cudax/stf: cuda_try migration — graph misc (PR6) by @andralex in https://github.com/NVIDIA/cccl/pull/9301
- [libcu++] Use == and < in cuda::args instead of <= by @pciolkosz in https://github.com/NVIDIA/cccl/pull/9884
- cudax/stf: cuda_try migration — stream_ctx (PR8) by @andralex in https://github.com/NVIDIA/cccl/pull/9306
- [STF] Migrate __stf/utility/ from cuda_safe_call to cuda_try by @andralex in https://github.com/NVIDIA/cccl/pull/9150
- Fix some miscellaneous clang-tidy errors by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9999
- Remove nvcc arch flags from clangd by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9996
- [nvtarget] Add support for
sm_107by @davebayer in https://github.com/NVIDIA/cccl/pull/10008 - [STF] Support compound while-loop conditions in the C API by @caugonnet in https://github.com/NVIDIA/cccl/pull/10006
- Small CUB doc corrections by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10011
- [libcu++] Optimize
cuda::std::rotlandcuda::std::rotrby @davebayer in https://github.com/NVIDIA/cccl/pull/10004 - Update
__cccl_ptx_isafor clang-cuda 22 by @davebayer in https://github.com/NVIDIA/cccl/pull/10007 - Fix unused comparison operators now that cuda::arguments only uses
operator<by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10021 - [cuda.compute]: Serialize hostjit builds and improve tests by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9998
- [cuda.compute]: Fix segmented sort selector race by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10024
- Use if in _CCCL_TRY_CUDA_API by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9985
- Restructure the RLE encode tuning policy by @nanan-nvidia in https://github.com/NVIDIA/cccl/pull/10028
- Enable
__float128incuda::std::fpclassify,isnormal,ilogbandlogbby @temujinkz in https://github.com/NVIDIA/cccl/pull/10014 - Flatten CUB/Thrust API docs by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10012
- Enable
__float128incuda::std::copysignandcuda::std::signbitby @temujinkz in https://github.com/NVIDIA/cccl/pull/9991 - Docs follow-up nits for flattened CUB/Thrust docs (#10012) by @gonidelis in https://github.com/NVIDIA/cccl/pull/10042
- [cuda.compute]: Add pytest-run-parallel by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9886
cuda_errorfixups by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10005- [libcu++] Make
DoNotOptimizefunctional on device by @davebayer in https://github.com/NVIDIA/cccl/pull/10037 - Extend ChainedPolicy test for
sm_107by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10047 - [STF] Expose geometry-aware allocation and native partition functions in the C API by @caugonnet in https://github.com/NVIDIA/cccl/pull/10038
- [libcu++] Implement P3793R2 Better Shifting (without SIMD) by @davebayer in #9993
- Exposes
cuda::execution::guaranteeby @elstehle in https://github.com/NVIDIA/cccl/pull/10022 - [cuda.compute]: Relax synchronization around Clang compilation in v2, keeping it only for linking by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10051
- [cuda.compute]: Fix windows race related to get nvrtc type name by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10049
- Remove trailing whitespace from CODEOWNERS by @pauleonix in https://github.com/NVIDIA/cccl/pull/10062
- [cuda.compute]: Fail Windows Python test jobs on native command errors by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10046
- [libcu++] Remove
cuda::__(shl|shr)by @davebayer in https://github.com/NVIDIA/cccl/pull/10050 - Minor QOL improvements to devcontainer launching script by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10055
- bugprone-exception-escape by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9510
- Prepare three-way partition tuning policies for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10077
- Prepare non-trivial-runs tuning policy for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10080
- [cuda.compute]: Add smoke tests for benchmarks by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9885
- Add pretty printers for cuda::buffer by @Jacobfaib in #9866
- Prepare scan-by-key tuning policy for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10079
- [cuda.compute]: bind nvJitLink to _12_0 aliases so cu12 wheels load on CTK 12.0 by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10083
- support for p0 explicit benchmarks by @srinivasyadav18 in https://github.com/NVIDIA/cccl/pull/10054
- Prepare reduce-by-key tuning policy for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10078
- Use block-scoped atomics in BlockHistogramAtomic on SM 60+ by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10084
- Take CUB environments by
const&by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10105 - [CUB] Fix DeviceAdjacentDifference for cuda::device_buffer iterators by @Noperi0r in #9861
- Implement prefetching by @gonidelis in https://github.com/NVIDIA/cccl/pull/9723
- [cudax][cuco] Reject undersized hyperloglog_ref storage by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/10220
- Enable
__float128incuda::stdcomparison functions (isgreaterfamily) by @temujinkz in https://github.com/NVIDIA/cccl/pull/10226 - [libcu++] Optimize
cuda::std::saturating_castby @davebayer in https://github.com/NVIDIA/cccl/pull/9724 - [cudax][cuco] Fix const mismatch in cooperative HyperLogLog merge by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/10217
- Split cuopt RAPIDS build by @bdice in https://github.com/NVIDIA/cccl/pull/10215
- Remove cuda.coop._experimental from cuda-cccl by @tpn in https://github.com/NVIDIA/cccl/pull/10104
- [cudax][cuco] Return double from HyperLogLog estimate by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/10218
- [libcu++] Fix
cuda::std::aligned_allocargument order by @davebayer in #9728 - Pin cuda-toolkit wheel to container's CTK major.minor in CI by @leofang in https://github.com/NVIDIA/cccl/pull/8160
- [cudax][cuco] Migrate host and device find APIs for fixed_capacity_map by @PointKernel in https://github.com/NVIDIA/cccl/pull/9868
- [cuda.compute]: Add ci matrix entry for minimal ft testing on windows by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9887
- Improve memory resource documentation by @bdice in https://github.com/NVIDIA/cccl/pull/10228
- [cuda.compute]: add CI job that builds c.parallel with thread sanitizer by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9986
- [libcudacxx][fp] add fpemu (double-precision emulation) + unit tests by @akolesov-nvidia in https://github.com/NVIDIA/cccl/pull/9777
- Implement
cuda::std::is_virtual_base_ofby @davebayer in #4397 - [libcu++] Avoid compilation issue with tuple_of_iterator_references by @miscco in https://github.com/NVIDIA/cccl/pull/10181
- [cuda.compute]: add back CI for CTK 12.0 by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10057
- [libcu++] Use
lit-style tests forfpemu(1/2) by @davebayer in https://github.com/NVIDIA/cccl/pull/10233 - Prepare select and partition tuning policies for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10076
- Prepare batched-copy tuning policy for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10081
- feat(security): onboard pre-commit and Pulse secret scanning by @gmanal in https://github.com/NVIDIA/cccl/pull/10010
- fix(ci): restore actions: read for secret-scan reusable workflow by @gmanal in https://github.com/NVIDIA/cccl/pull/10515
- Implement
P2835R7andP3936R1std::atomic_ref::address()by @charan-003 in #9621 - cudax/stf: cuda_try migration — stream_task (PR9) by @andralex in https://github.com/NVIDIA/cccl/pull/10003
- Prevent OOB write for RBK and RLE in streaming context by @nanan-nvidia in #10504
- Replace deprecated
std::aligned_storage_tin STFsmall_vectorby @caugonnet in https://github.com/NVIDIA/cccl/pull/10507 - Add pretty printers for cuda::std::array by @pieroevcc in #10528
- [libcu++] Avoid use of
__CUDA_ARCH__in fpemu by @davebayer in https://github.com/NVIDIA/cccl/pull/10385 - [libcu++] Fix
_Float64test for fpemu in C++23 by @davebayer in https://github.com/NVIDIA/cccl/pull/10508 - [cuco] Adds support for 1 and 2 byte key and value types in
fixed_capacity_mapby @ryanjspears in https://github.com/NVIDIA/cccl/pull/10025 - cudax/stf: cuda_try migration — graph_task (PR10) by @andralex in https://github.com/NVIDIA/cccl/pull/10523
- [libcu++] Implement internal resizable buffer by @pciolkosz in https://github.com/NVIDIA/cccl/pull/10221
- Add NVRTC compatibility errors to CUB headers by @hzaidi05 in #7079
- Reduce PR NVHPC CI coverage by @jrhemstad in https://github.com/NVIDIA/cccl/pull/10527
- Restore _CCCL_GRID_CONSTANT on streaming_context in DeviceSelect by @nanan-nvidia in https://github.com/NVIDIA/cccl/pull/10537
- [Infra] Split SM75 GPU runners more evenly between RTX2080 and T4 by @miscco in https://github.com/NVIDIA/cccl/pull/10513
- Add environment docs landing page with essential info by @gonidelis in https://github.com/NVIDIA/cccl/pull/10043
- [libcu++] Do not instantiate types for inline variables of constructible traits by @miscco in https://github.com/NVIDIA/cccl/pull/10250
- [libcu++] Implement P3104R6 Bit permutations by @davebayer in #10063
- [Tile] Mark formatters of tuning policies as host_device by @miscco in https://github.com/NVIDIA/cccl/pull/10539
- [Tile] Wrap
char_traits::eqin a functor by @miscco in https://github.com/NVIDIA/cccl/pull/10542 - [cudax] Exempt places CMake files from cmake-codeowners review by @caugonnet in https://github.com/NVIDIA/cccl/pull/9782
- [Tile] Mark fpemu as unsupported in tile mode by @miscco in https://github.com/NVIDIA/cccl/pull/10538
- [libcu++] Fix
cuda::make_tma_descriptor(...)by @davebayer in https://github.com/NVIDIA/cccl/pull/10546 - [Tile] Mark all atomics functions as host device only by @miscco in https://github.com/NVIDIA/cccl/pull/10543
- Fixup _CCCL_HOST_API and inline constexpr variables in memory land by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10536
- Fix more static constexpr -> inline constexpr by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10549
- CUB: Add a custom test wrapper with memory classification by @ratnachandugembali-lgtm in https://github.com/NVIDIA/cccl/pull/10529
- [Tile] Fix complex interop with extended floating point types by @miscco in https://github.com/NVIDIA/cccl/pull/10550
- [TILE] Add a workaround to make
invokework in tile mode by @miscco in https://github.com/NVIDIA/cccl/pull/10545 - cudax/stf: add defer_exception and void checks for SCOPE by @andralex in https://github.com/NVIDIA/cccl/pull/10559
- [libcudacxx] Add GDB/LLDB pretty-printers for cuda::std::complex and cuda::complex by @HenrikGharagyozyan in #10563
- Add tunable prefetching to DeviceSelect::Flagged by @anikaj-eng in https://github.com/NVIDIA/cccl/pull/10519
- [cub] Always set smem limit to max for lookahead scan by @davebayer in https://github.com/NVIDIA/cccl/pull/10570
- Use devcontainers from RAPIDS 26.10 by @bdice in https://github.com/NVIDIA/cccl/pull/10526
- [libcu++] Add new memory pool attributes from CUDA 13.3 by @pciolkosz in #9798
- cudax: complete Library Fundamentals TS v3 scope guards by @andralex in https://github.com/NVIDIA/cccl/pull/10565
- Compile time benchmarking tool by @griwes in https://github.com/NVIDIA/cccl/pull/9498
- [Backport branch/3.5.x] [libcu++] Allow no GPUDirect RDMA flush options by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/10633
- [Backport branch/3.5.x] Fixes a race condition in
DeviceTopKby @github-actions[bot] in #10683 - [Backport branch/3.5.x] [Tile] Disable tile support by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/10969
- [Backport branch/3.5.x] [libcu++] Fix the base 10 width of exact powers of ten in to_chars by @github-actions[bot] in #10990
- [Backport branch/3.5.x] [CUB] Fix unqualified calls to libcu++ entities by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11015
- [backport 3.5] Remove fpemu from 3.5 release by @davebayer in https://github.com/NVIDIA/cccl/pull/10613
- [Backport branch/3.5.x] [libcu++] Allow alternate pinned memory type reporting by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/10632
- [backport 3.5.x] Remove
constant_wrapperfrom 3.5 release by @davebayer in https://github.com/NVIDIA/cccl/pull/11065 - [Backport branch/3.5.x] [libcu++] Add
TREAT_WARNINGS_AS_ERRORS.litparser by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11087 - [Backport branch/3.5.x] [libcu++] Disable flaky test by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11114
- [Backport to 3.5] Fix MSVC26 (#11103) by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/11110
- [Backport 3.5] Update cuda::ptx for CUDA 13.4 by @pciolkosz in #11054
- [Backport branch/3.5.x] Add inputs for controlling which repo and branch are checked out by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11009
- [Backport branch/3.5.x] Deduplicate and extend nightly/weekly workflows. (#11074) by @wmaxey in https://github.com/NVIDIA/cccl/pull/11162
- [Backport branch/3.5.x] [CI] Add fields in CI matrix that determine devcontainer repo and runner labels by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11125
- [Backport 3.5] Backport #10747 and #11111 by @miscco in https://github.com/NVIDIA/cccl/pull/11201
- [Backport branch/3.5.x] Fixup - remove extra parameter in workflow call by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11194
- [Backport branch/3.5.x] [libcu++] Fix
__builtin_bswapgnot being supported by nvcc by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11178 - [Backport 3.5] Backport
tupleandpairimprovements by @miscco in https://github.com/NVIDIA/cccl/pull/11475 - [Backport branch/3.5.x] [libcu++] Fix
arch_traitsfor sm100 by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11484
New Contributors
- @vip892766gma made their first contribution in #8980
- @arnavnagzirkar made their first contribution in #9211
- @Oxygen56 made their first contribution in #9196
- @AneeshGidda made their first contribution in #8780
- @nethum529 made their first contribution in #9680
- @Functionhx made their first contribution in #9785
- @codingwithmagga made their first contribution in https://github.com/NVIDIA/cccl/pull/9383
- @rkothari3 made their first contribution in https://github.com/NVIDIA/cccl/pull/9833
- @Noperi0r made their first contribution in #9861
- @gmanal made their first contribution in https://github.com/NVIDIA/cccl/pull/10010
- @pieroevcc made their first contribution in #10528
- @hzaidi05 made their first contribution in #7079
- @anikaj-eng made their first contribution in https://github.com/NVIDIA/cccl/pull/10519
Full Changelog: v3.5.0.dev...v3.5.0