Repository navigation
v1.18.0 #2035
shi-eric
announced in
Announcements
v1.18.0
#2035
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Warp v1.18.0
Warp v1.18 introduces the
warp.geometrymodule, with Python and kernel APIs for extracting meshes from implicit fields and sampled rigid motion, and improving 2D triangulations directly on the device. Experimental CPU block execution improves tile portability by supporting the same explicitblock_dimon CPU and CUDA. The release also supports automatic differentiation through vector and matrix component updates in arrays and avoids unnecessary GPU-to-CPU transfers of sparse-matrix block counts.Important
The default PyPI and nightly wheels for Linux and Windows now use CUDA Toolkit 13.4. CUDA acceleration requires an NVIDIA R580-series or newer driver and a Turing (
sm_75) or newer GPU. CUDA 12 compatibility wheels remain available from GitHub Releases with a+cu12version suffix.New features
Sparse marching cubes
You can now extract an isosurface from a custom sparse volume or other geometry representation without converting it to a dense grid.
warp.geometry.sparse_marching_cubes()accepts a callable that samples a scalar field at batches of query points, so your sampler can read directly from the existing data structure. Warp builds a Lipschitz octree to prune cells away from the requested surface and runs marching cubes on the remaining cells (#1803).nx,ny, andnzspecify grid-node counts, as in dense marching cubes. The field callable acceptswp.vec3query points in batches and returnswp.float32values on the same device. The output containswp.vec3vertices and flat triangle indices, with three indices per face. This example evaluates a sphere field in a Warp kernel on the GPU and compares its evaluation count with sampling every node of a257³grid:Output:
Here, sparse extraction uses about 2.4% of the field evaluations required by dense sampling.
You can differentiate extracted vertex positions with respect to sampled field values. Allocate the field output with
requires_grad=Trueand extract insidewp.Tape(). Cell selection and mesh connectivity remain fixed during differentiation.The default
lipschitz_bound=1.0works for a true signed distance function. Fields that change faster with distance need a larger bound. Setting it too low can discard cells containing the surface.Parts of the surface outside
lowerandupperare omitted, so the mesh may be open at those bounds. See the sparse marching cubes example for a mesh-query field and a comparison with dense extraction.If you already have a list of grid cells and field values at their eight corners, use
warp.geometry.sparse_marching_cubes_from_cells(). It extracts the mesh from those cells without building an octree.Delaunay edge flipping
warp.geometry.delaunay_edge_flip()updates a 2D triangle mesh's connectivity on the device after its vertices move. It flips non-Delaunay interior edges in place.warp.geometry.tri_tri_adjacency()identifies each edge's neighboring triangle and the matching edge index within that neighbor (#1797).The shared diagonal changes from
(0, 1)to(2, 3).Edge flipping requires a manifold mesh whose triangle vertices are consistently ordered counterclockwise. It stops when no more flips are needed or
max_passesis reached. If you supply reference positions, the algorithm skips flips that would create degenerate triangles in the reference mesh. Both APIs support CUDA graph capture. Supplyvertex_countto the adjacency helper during capture, and read the flip count after replay.The elastic shape optimization example uses device-side edge flipping while optimizing a deforming mesh.
Swept-volume meshing
warp.geometry.swept_volume_mesh()builds a single mesh of the space occupied by a rigid assembly across sampled poses. It samples a signed distance field on a dense grid. Each grid point queries every mesh at every supplied pose, so finer grids and more pose samples increase computation. Use the result for motion visualization or as an input to robotics clearance checks (#1824).The sampled tetrahedron spans
(0, 0, 0)to(2, 1, 1); the positive threshold expands the mesh beyond those bounds.Transforms have shape
(mesh_count, sample_count, 7), with translation followed by anxyzwquaternion. The default sign mode requires each input mesh to be built withsupport_winding_number=True. The newmesh.support_winding_numberproperty lets you check that requirement.The envelope covers the supplied poses. It does not conservatively bound motion between samples, so fast or thin objects need sufficiently dense pose sampling. To enclose the sampled poses despite grid discretization, use a positive
thresholdgreater than half the grid-cell diagonal. The example uses0.09with a cubic cell size of0.1.See the swept-volume example for a procedural robot arm, animated USD input, and alternative sign modes.
Dense isosurface extraction
Dense isosurface extraction now has a common
warp.geometry.IsoSurfaceBaseinterface, implemented bywarp.geometry.IsoSurfaceMarchingCubes. Callextract()on the class to return vertices and triangle indices without creating an extractor instance (#1614):The input must be a three-dimensional
wp.float32array with at least two nodes on every axis. You can differentiate the extracted vertex positions with respect to field values, with mesh connectivity held fixed. The isosurface example has moved intowarp/examples/geometry/.Tile programming
Cooperative CPU block execution
Important
This is an experimental feature. The API may change without a formal deprecation cycle.
Tile kernels that require multiple cooperating logical threads within a block can now produce matching results on CPU and CUDA using the same explicit
block_dim. Setwp.config.enable_cpu_blocks = Trueto enable CPU blocks with 2 to 1024 logical kernel threads. Each logical thread has its own index within the block. Threads in the same block can synchronize and share tile data (#1638).CPU execution defaults to one logical kernel thread per block. Cooperative CPU blocks schedule their logical threads on a single operating-system thread and can be substantially slower. They do not add CPU parallelism or vectorize execution across those logical threads. AddressSanitizer builds reject enabled CPU launches with
block_dim > 1. See CPU tile execution for portability details.Differentiation and function control
Automatic differentiation through array component updates
Automatic differentiation now supports assignments and updates to individual components of array elements, such as vectors and matrices. Supported patterns include
out[i].x = src[i],out[i][j] += value, andmatrices[i][r, c] = value(#1451).You no longer need to load the whole array element, modify a local copy, and write it back to differentiate a supported component assignment.
Function inlining
@wp.funcnow accepts aninlineoption. Inlining removes call overhead by inserting the function body at each call site, which can increase code size. Set@wp.func(inline=True)to inline the body,inline=Falseto keep a separate function, orinline=Noneto let the compiler decide. The choice also applies to generated backward functions (#1849).Sparse matrix workflows
On-demand sparse matrix block counts
Sparse matrix construction and structural updates now avoid immediately copying the stored block count from the GPU to the CPU. After compact construction,
matrix.nnzcan still be an upper bound. For a compact matrix, callmatrix.nnz_sync()when you need the exact number of stored matrix blocks, including before slicing the column-index or value arrays (#1792).Replace the deprecated asynchronous block-count transfer with
nnz_sync()when you need the count on the CPU:Call
nnz_sync()outside CUDA graph capture. The deprecated asynchronous method remains available, but its pending count transfer can be unsafe when matrix updates use multiple streams or when the transfer is captured and replayed.Performance improvements
Kernels that enumerate collision candidates with axis-aligned bounding-box (AABB) queries on
wp.Bvhandwp.Meshran 1.1 to 2.1× as fast as Warp 1.17 in a synthetic benchmark. The benchmark used 200,000 queries over a 122,018-triangle heightfield,lbvhandsahtree constructors with leaf-size settings of 1 and 8 triangles, and an RTX PRO 6000 Blackwell Server Edition MIG 1g.24gb GPU partition (#1840, #1843).wp.tile_matmul()used 29% less kernel time than Warp 1.17 for 1024×1024 FP32 matrix multiplication with 64×64×64 tiles andblock_dim=256, on the same GPU partition used for the AABB benchmark (#1938). Gains depend on the matrix and tile configuration.New and updated examples
core/example_marching_cubes.pytogeometry/example_isosurface.pyand now useswarp.geometry.IsoSurfaceMarchingCubes, which implements thewarp.geometry.IsoSurfaceBaseinterface (Add wp.DualContouring and wp.IsoSurfaceBase interface for swappable isosurface extraction #1614).Breaking changes
32-bit integer texture sampling is now rejected
Sampling
wp.uint32andwp.int32textures is now rejected by a CPU process abort or CUDA kernel trap. These types remain available for storage, copies, and interop. To reproduce the previous CPU sampling behavior, convert data towp.float32and normalize unsigned values to[0, 1]or signed values to[-1, 1]before creating the texture (#1731).For unsigned data, where
rawis a NumPyuint32array:Invalid inputs now raise different exception types
Invalid inputs in the JAX foreign function interface (FFI), optimizers, DLPack, OpenGL rendering, and
warp.fem.PointBasisSpacenow raise more specific exception types. Update handlers that relied on the oldAssertionErrororValueErrorclasses. For example, SGD gradient dtype mismatches now raiseTypeError, while gradient shape mismatches raiseValueError(#1930).Bug fixes
Signed integer floor division now matches Python
//andwp.floordiv()now round signed integer quotients toward negative infinity, matching Python and NumPy (#1918). Review code or tests that relied on truncation toward zero:For nonnegative operands and a positive divisor, consider unsigned types such as
wp.uint32orwp.uint64: unsigned//avoids the signed rounding correction. Choose a type that covers the value range, and watch for intermediate subtraction that could go negative.Sparse expressions now scale the intended matrix
Expressions with a scaled sparse matrix on the left of
+or-now apply the scale to that matrix. Previously,(scale * A) + Band(scale * A) - Bincorrectly scaledB. The addition workaround of putting the scaled expression on the right is no longer needed (#1910):Tile range element counts now round up
wp.tile_arange()now rounds the element count up when the span is not divisible by the step (#1774). The integer rangewp.tile_arange(0, 10, 3)returns[0, 3, 6, 9], with the stop value10excluded. Update any output array allocated using the previous count:Announcements
Removals in this release
warp.jax_experimentalis removed. Use the top-level Warp JAX APIs. The legacy custom-call implementation and graph-cache default getter and setter are also removed. Passgraph_cache_maxdirectly towp.jax_callable()to configure that wrapper. Review legacy wrappers against the current JAX integration guide (Remove deprecated JAX compatibility APIs #1904).warp.config.verboseandwarp.config.quietare removed. Setwarp.config.log_levelto select debug or warning-level output (Remove deprecated logging configuration flags #1905).wp.HashGridQueryHandwp.HashGridQueryDare removed. Use the commonwp.HashGridQuerytype (Remove deprecated hash grid query aliases #1906).Upcoming removals
wp.MarchingCubesis deprecated. Explicitly importwarp.geometryand usewarp.geometry.IsoSurfaceMarchingCubes. The top-level name remains an alias during the standard deprecation period. We will remove it in a future Warp feature release. The deprecated domain-bound constructor arguments and attributes remain available as aliases oflowerandupper.*_tiledquery spellings are deprecated and targeted for removal in Warp 1.22. Replace them withwp.tile_bvh_query_aabb(),wp.tile_bvh_query_ray(),wp.tile_bvh_query_next(),wp.tile_mesh_query_aabb(), andwp.tile_mesh_query_aabb_next(). Query behavior is unchanged, and compiling the old names emits aDeprecationWarningidentifying the replacement (Deprecate*_tiledtile query aliases #1885).IsoSurfaceMarchingCubes.extract_surface_marching_cubes()is deprecated. Useextract()instead. The alias remains available during the standard deprecation period. We will remove it in a future Warp feature release (Add wp.DualContouring and wp.IsoSurfaceBase interface for swappable isosurface extraction #1614).Platform support
sm_75) or newer GPU. For CUDA 12 environments, choose a+cu12compatibility wheel from GitHub Releases, or build Warp from source with CUDA 12 (Publish PyPI wheels using CUDA 13.4 #1955).--no-use-libmathdxtobuild_lib.py. See the source-build guide for details.warp.lib,warp-clang.lib, and the newwarp_clang.hheader let native C++ applications link without resolving each entry point manually. Keep headers, import libraries, and DLLs on the same Warp release (Provide an import .lib for Warp's .dll(s) on Windows #1888).Acknowledgments
We also thank the following contributors from outside the core Warp development team:
//) truncates toward zero instead of flooring #1918).wp.utils.array_scan()raise a Python exception when a CUDA scan fails ([BUG]array_scansilently corrupts its output on device OOM: uncheckedwp_alloc_devicereturn launches CUB on a NULL buffer #1894).For a complete list of changes, see the full changelog.
This discussion was created from the release v1.18.0.
All reactions