Skip to content

v1.18.0

Latest

Choose a tag to compare

@github-actions github-actions released this 05 Oct 00:24
· 47 commits to main since this release
v1.18.0
f2eaed8

Warp v1.18.0

Warp v1.18 introduces the warp.geometry module, with Python and kernel APIs for extracting meshes from implicit fields and sampled rigid motion, and improving 2D triangulations directly on the device. Experimental CPU block execution improves tile portability by supporting the same explicit block_dim on CPU and CUDA. The release also supports automatic differentiation through vector and matrix component updates in arrays and avoids unnecessary GPU-to-CPU transfers of sparse-matrix block counts.

Important

The default PyPI and nightly wheels for Linux and Windows now use CUDA Toolkit 13.4. CUDA acceleration requires an NVIDIA R580-series or newer driver and a Turing (sm_75) or newer GPU. CUDA 12 compatibility wheels remain available from GitHub Releases with a +cu12 version suffix.

New features

Sparse marching cubes

You can now extract an isosurface from a custom sparse volume or other geometry representation without converting it to a dense grid. warp.geometry.sparse_marching_cubes() accepts a callable that samples a scalar field at batches of query points, so your sampler can read directly from the existing data structure. Warp builds a Lipschitz octree to prune cells away from the requested surface and runs marching cubes on the remaining cells (#1803).

nx, ny, and nz specify grid-node counts, as in dense marching cubes. The field callable accepts wp.vec3 query points in batches and returns wp.float32 values on the same device. The output contains wp.vec3 vertices and flat triangle indices, with three indices per face. This example evaluates a sphere field in a Warp kernel on the GPU and compares its evaluation count with sampling every node of a 257³ grid:

import warp as wp
import warp.geometry

@wp.kernel
def sample_sphere(points: wp.array[wp.vec3], values: wp.array[float]):
    i = wp.tid()
    values[i] = wp.length(points[i]) - 0.5

def field(points):
    values = wp.empty(points.shape, dtype=wp.float32, device=points.device)
    wp.launch(sample_sphere, dim=points.shape[0], inputs=[points],
              outputs=[values], device=points.device)
    return values

n = 257
vertices, indices, stats = warp.geometry.sparse_marching_cubes(
    field, n, n, n,
    lower=(-1.0, -1.0, -1.0), upper=(1.0, 1.0, 1.0),
    return_stats=True,
)
print(f"Mesh: {vertices.shape[0]:,} vertices, {indices.shape[0] // 3:,} triangles")
print(f"Sparse field evaluations: {stats['field_evaluations']:,}")
print(f"Dense grid samples: {n**3:,}")

Output:

Mesh: 77,094 vertices, 154,184 triangles
Sparse field evaluations: 400,259
Dense grid samples: 16,974,593

Here, sparse extraction uses about 2.4% of the field evaluations required by dense sampling.

You can differentiate extracted vertex positions with respect to sampled field values. Allocate the field output with requires_grad=True and extract inside wp.Tape(). Cell selection and mesh connectivity remain fixed during differentiation.

The default lipschitz_bound=1.0 works for a true signed distance function. Fields that change faster with distance need a larger bound. Setting it too low can discard cells containing the surface.

Parts of the surface outside lower and upper are omitted, so the mesh may be open at those bounds. See the sparse marching cubes example for a mesh-query field and a comparison with dense extraction.

If you already have a list of grid cells and field values at their eight corners, use warp.geometry.sparse_marching_cubes_from_cells(). It extracts the mesh from those cells without building an octree.

Delaunay edge flipping

warp.geometry.delaunay_edge_flip() updates a 2D triangle mesh's connectivity on the device after its vertices move. It flips non-Delaunay interior edges in place. warp.geometry.tri_tri_adjacency() identifies each edge's neighboring triangle and the matching edge index within that neighbor (#1797).

import warp as wp
import warp.geometry

positions = wp.array([[-3.0, 0.0], [3.0, 0.0], [0.0, 1.0], [0.0, -1.0]],
                     dtype=wp.vec2)
triangles = wp.array([[0, 1, 2], [1, 0, 3]], dtype=wp.int32)
flips = warp.geometry.delaunay_edge_flip(positions, triangles)
neighbors, neighbor_edges = warp.geometry.tri_tri_adjacency(triangles, vertex_count=4)
print(int(flips.numpy()[0]))  # 1
faces = triangles.numpy().tolist()
print(sorted(set(faces[0]) & set(faces[1])))  # [2, 3]

The shared diagonal changes from (0, 1) to (2, 3).

Edge flipping requires a manifold mesh whose triangle vertices are consistently ordered counterclockwise. It stops when no more flips are needed or max_passes is reached. If you supply reference positions, the algorithm skips flips that would create degenerate triangles in the reference mesh. Both APIs support CUDA graph capture. Supply vertex_count to the adjacency helper during capture, and read the flip count after replay.

The elastic shape optimization example uses device-side edge flipping while optimizing a deforming mesh.

Swept-volume meshing

warp.geometry.swept_volume_mesh() builds a single mesh of the space occupied by a rigid assembly across sampled poses. It samples a signed distance field on a dense grid. Each grid point queries every mesh at every supplied pose, so finer grids and more pose samples increase computation. Use the result for motion visualization or as an input to robotics clearance checks (#1824).

import numpy as np
import warp as wp
import warp.geometry

points = wp.array([[0, 0, 0], [1, 0, 0], [0, 1, 0], [0, 0, 1]], dtype=wp.vec3)
faces = wp.array([0, 2, 1, 0, 3, 2, 0, 1, 3, 1, 2, 3], dtype=wp.int32)
mesh = wp.Mesh(points, faces, support_winding_number=True)
poses = np.zeros((1, 8, 7), dtype=np.float32)
poses[..., 6] = 1.0  # Identity rotations, quaternion order xyzw.
poses[0, :, 0] = np.linspace(0.0, 1.0, 8)
vertices, indices = warp.geometry.swept_volume_mesh(
    [mesh], poses, voxel_size=0.1, threshold=0.09
)
v = vertices.numpy()
print(np.round(v.min(axis=0), 2))  # [-0.09 -0.09 -0.09]
print(np.round(v.max(axis=0), 2))  # [2.09 1.09 1.09]

The sampled tetrahedron spans (0, 0, 0) to (2, 1, 1); the positive threshold expands the mesh beyond those bounds.

Transforms have shape (mesh_count, sample_count, 7), with translation followed by an xyzw quaternion. The default sign mode requires each input mesh to be built with support_winding_number=True. The new mesh.support_winding_number property lets you check that requirement.

The envelope covers the supplied poses. It does not conservatively bound motion between samples, so fast or thin objects need sufficiently dense pose sampling. To enclose the sampled poses despite grid discretization, use a positive threshold greater than half the grid-cell diagonal. The example uses 0.09 with a cubic cell size of 0.1.

See the swept-volume example for a procedural robot arm, animated USD input, and alternative sign modes.

Dense isosurface extraction

Dense isosurface extraction now has a common warp.geometry.IsoSurfaceBase interface, implemented by warp.geometry.IsoSurfaceMarchingCubes. Call extract() on the class to return vertices and triangle indices without creating an extractor instance (#1614):

import numpy as np
import warp as wp
import warp.geometry

x, y, z = np.meshgrid(*([np.linspace(-1.0, 1.0, 5)] * 3), indexing="ij")
field = wp.array(np.sqrt(x*x + y*y + z*z) - 0.6, dtype=wp.float32)
vertices, indices = warp.geometry.IsoSurfaceMarchingCubes.extract(
    field, lower=(-1.0, -1.0, -1.0), upper=(1.0, 1.0, 1.0)
)
print(vertices.shape[0], indices.shape[0] // 3)  # 30 56

The input must be a three-dimensional wp.float32 array with at least two nodes on every axis. You can differentiate the extracted vertex positions with respect to field values, with mesh connectivity held fixed. The isosurface example has moved into warp/examples/geometry/.

Tile programming

Cooperative CPU block execution

Important

This is an experimental feature. The API may change without a formal deprecation cycle.

Tile kernels that require multiple cooperating logical threads within a block can now produce matching results on CPU and CUDA using the same explicit block_dim. Set wp.config.enable_cpu_blocks = True to enable CPU blocks with 2 to 1024 logical kernel threads. Each logical thread has its own index within the block. Threads in the same block can synchronize and share tile data (#1638).

import warp as wp

wp.config.enable_cpu_blocks = True

@wp.kernel
def sum_thread_indices(out: wp.array[float]):
    block_index, thread_index = wp.tid()
    values = wp.tile(float(thread_index))
    wp.tile_store(out, wp.tile_sum(values))

out = wp.empty(1, dtype=wp.float32, device="cpu")
wp.launch_tiled(sum_thread_indices, dim=1, outputs=[out], block_dim=4, device="cpu")
print(out.numpy())  # [6.]

CPU execution defaults to one logical kernel thread per block. Cooperative CPU blocks schedule their logical threads on a single operating-system thread and can be substantially slower. They do not add CPU parallelism or vectorize execution across those logical threads. AddressSanitizer builds reject enabled CPU launches with block_dim > 1. See CPU tile execution for portability details.

Differentiation and function control

Automatic differentiation through array component updates

Automatic differentiation now supports assignments and updates to individual components of array elements, such as vectors and matrices. Supported patterns include out[i].x = src[i], out[i][j] += value, and matrices[i][r, c] = value (#1451).

import warp as wp

@wp.kernel
def write_component(src: wp.array[float], out: wp.array[wp.vec3]):
    i = wp.tid()
    out[i].x += 2.0 * src[i]

src = wp.array([3.0], dtype=wp.float32, requires_grad=True)
out = wp.zeros(1, dtype=wp.vec3, requires_grad=True)
with wp.Tape() as tape:
    wp.launch(write_component, dim=1, inputs=[src], outputs=[out])
tape.backward(grads={out: wp.ones_like(out)})
print(src.grad.numpy())  # [2.]

You no longer need to load the whole array element, modify a local copy, and write it back to differentiate a supported component assignment.

Function inlining

@wp.func now accepts an inline option. Inlining removes call overhead by inserting the function body at each call site, which can increase code size. Set @wp.func(inline=True) to inline the body, inline=False to keep a separate function, or inline=None to let the compiler decide. The choice also applies to generated backward functions (#1849).

Sparse matrix workflows

On-demand sparse matrix block counts

Sparse matrix construction and structural updates now avoid immediately copying the stored block count from the GPU to the CPU. After compact construction, matrix.nnz can still be an upper bound. For a compact matrix, call matrix.nnz_sync() when you need the exact number of stored matrix blocks, including before slicing the column-index or value arrays (#1792).

import warp as wp
import warp.sparse as sparse

rows = wp.array([0, 0, 1], dtype=wp.int32)
columns = wp.array([0, 0, 1], dtype=wp.int32)
values = wp.array([1.0, 2.0, 3.0], dtype=wp.float32)
matrix = sparse.bsr_from_triplets(2, 2, rows, columns, values)
print(matrix.nnz)  # 3: upper bound before synchronization
exact_count = matrix.nnz_sync()
print(exact_count)  # 2

Replace the deprecated asynchronous block-count transfer with nnz_sync() when you need the count on the CPU:

- matrix.copy_nnz_async()
+ exact_count = matrix.nnz_sync()

Call nnz_sync() outside CUDA graph capture. The deprecated asynchronous method remains available, but its pending count transfer can be unsafe when matrix updates use multiple streams or when the transfer is captured and replayed.

Performance improvements

Kernels that enumerate collision candidates with axis-aligned bounding-box (AABB) queries on wp.Bvh and wp.Mesh ran 1.1 to 2.1× as fast as Warp 1.17 in a synthetic benchmark. The benchmark used 200,000 queries over a 122,018-triangle heightfield, lbvh and sah tree constructors with leaf-size settings of 1 and 8 triangles, and an RTX PRO 6000 Blackwell Server Edition MIG 1g.24gb GPU partition (#1840, #1843).

wp.tile_matmul() used 29% less kernel time than Warp 1.17 for 1024×1024 FP32 matrix multiplication with 64×64×64 tiles and block_dim=256, on the same GPU partition used for the AABB benchmark (#1938). Gains depend on the matrix and tile configuration.

New and updated examples

  • The sparse marching cubes example remeshes the Stanford bunny from a signed distance field. It includes a dense extraction mode and an option to display the octree leaf cells. A companion benchmark compares sparse and dense extraction for analytic and mesh-query fields (#1803).
  • The swept volume example computes the motion envelope from sampled poses of a two-link arm or an animated USD assembly (#1824).
  • The dense isosurface example has moved from core/example_marching_cubes.py to geometry/example_isosurface.py and now uses warp.geometry.IsoSurfaceMarchingCubes, which implements the warp.geometry.IsoSurfaceBase interface (#1614).
  • The cantilever topology optimization example optimizes a cantilever's material distribution using double-precision FEM and NLopt MMA. It filters densities, computes derivatives with automatic differentiation, and displays the density during optimization (#1922).
  • The fluid checkpointing example reduces memory use during differentiation with a custom backward pass that reuses scratch grids instead of storing every intermediate Jacobi state. The original and custom-backward fluid checkpointing examples now size checkpoint segments from estimated grid storage, report current and peak CUDA memory use after the first optimization iteration, and use compatible discretizations for divergence and the pressure gradient (#1963).

Breaking changes

32-bit integer texture sampling is now rejected

Sampling wp.uint32 and wp.int32 textures is now rejected by a CPU process abort or CUDA kernel trap. These types remain available for storage, copies, and interop. To reproduce the previous CPU sampling behavior, convert data to wp.float32 and normalize unsigned values to [0, 1] or signed values to [-1, 1] before creating the texture (#1731).

For unsigned data, where raw is a NumPy uint32 array:

- texture = wp.Texture2D(data=raw)
+ normalized = raw.astype(np.float32) / np.float32(np.iinfo(np.uint32).max)
+ texture = wp.Texture2D(data=normalized)

Invalid inputs now raise different exception types

Invalid inputs in the JAX foreign function interface (FFI), optimizers, DLPack, OpenGL rendering, and warp.fem.PointBasisSpace now raise more specific exception types. Update handlers that relied on the old AssertionError or ValueError classes. For example, SGD gradient dtype mismatches now raise TypeError, while gradient shape mismatches raise ValueError (#1930).

Bug fixes

Signed integer floor division now matches Python

// and wp.floordiv() now round signed integer quotients toward negative infinity, matching Python and NumPy (#1918). Review code or tests that relied on truncation toward zero:

- expected = [-2]  # Warp 1.17: -7 // 3 truncated toward zero.
+ expected = [-3]  # Warp 1.18: -7 // 3 rounds down.
import warp as wp

@wp.kernel
def divide(a: int, b: int, out: wp.array[int]):
    out[0] = a // b

out = wp.empty(1, dtype=wp.int32)
wp.launch(divide, dim=1, inputs=[-7, 3], outputs=[out])
print(out.numpy())  # [-3]

For nonnegative operands and a positive divisor, consider unsigned types such as wp.uint32 or wp.uint64: unsigned // avoids the signed rounding correction. Choose a type that covers the value range, and watch for intermediate subtraction that could go negative.

Sparse expressions now scale the intended matrix

Expressions with a scaled sparse matrix on the left of + or - now apply the scale to that matrix. Previously, (scale * A) + B and (scale * A) - B incorrectly scaled B. The addition workaround of putting the scaled expression on the right is no longer needed (#1910):

- C = sparse.bsr_copy(B + (2.0 * A))  # Warp 1.17 workaround.
+ C = sparse.bsr_copy((2.0 * A) + B)
import warp as wp
import warp.sparse as sparse

indices = wp.array([0, 1], dtype=wp.int32)
A = sparse.bsr_from_triplets(2, 2, indices, indices,
                             wp.array([1.0, 2.0], dtype=wp.float32))
B = sparse.bsr_from_triplets(2, 2, indices, indices,
                             wp.array([3.0, 4.0], dtype=wp.float32))
C = sparse.bsr_copy((2.0 * A) + B)
print(C.values.numpy()[:C.nnz_sync()])  # [5. 8.]

Tile range element counts now round up

wp.tile_arange() now rounds the element count up when the span is not divisible by the step (#1774). The integer range wp.tile_arange(0, 10, 3) returns [0, 3, 6, 9], with the stop value 10 excluded. Update any output array allocated using the previous count:

- out = wp.empty(3, dtype=wp.int32)
+ out = wp.empty(4, dtype=wp.int32)
import warp as wp

@wp.kernel
def make_range(out: wp.array[int]):
    wp.tile_store(out, wp.tile_arange(0, 10, 3, dtype=wp.int32))

out = wp.empty(4, dtype=wp.int32)
wp.launch_tiled(make_range, dim=1, outputs=[out], block_dim=4)
print(out.numpy())  # [0 3 6 9]

Announcements

Removals in this release

  • warp.jax_experimental is removed. Use the top-level Warp JAX APIs. The legacy custom-call implementation and graph-cache default getter and setter are also removed. Pass graph_cache_max directly to wp.jax_callable() to configure that wrapper. Review legacy wrappers against the current JAX integration guide (#1904).
- from warp.jax_experimental import jax_kernel
+ from warp import jax_kernel
  • warp.config.verbose and warp.config.quiet are removed. Set warp.config.log_level to select debug or warning-level output (#1905).
- wp.config.verbose = True
+ wp.config.log_level = wp.LOG_DEBUG
- wp.config.quiet = True
+ wp.config.log_level = wp.LOG_WARNING
  • wp.HashGridQueryH and wp.HashGridQueryD are removed. Use the common wp.HashGridQuery type (#1906).
- query: wp.HashGridQueryH
+ query: wp.HashGridQuery

Upcoming removals

  • wp.MarchingCubes is deprecated. Explicitly import warp.geometry and use warp.geometry.IsoSurfaceMarchingCubes. The top-level name remains an alias during the standard deprecation period. We will remove it in a future Warp feature release. The deprecated domain-bound constructor arguments and attributes remain available as aliases of lower and upper.
  • The *_tiled query spellings are deprecated and targeted for removal in Warp 1.22. Replace them with wp.tile_bvh_query_aabb(), wp.tile_bvh_query_ray(), wp.tile_bvh_query_next(), wp.tile_mesh_query_aabb(), and wp.tile_mesh_query_aabb_next(). Query behavior is unchanged, and compiling the old names emits a DeprecationWarning identifying the replacement (#1885).
  • IsoSurfaceMarchingCubes.extract_surface_marching_cubes() is deprecated. Use extract() instead. The alias remains available during the standard deprecation period. We will remove it in a future Warp feature release (#1614).

Platform support

  • The default Linux and Windows wheels switch to CUDA Toolkit 13.4. For CUDA acceleration, use an NVIDIA R580-series or newer driver and a Turing (sm_75) or newer GPU. For CUDA 12 environments, choose a +cu12 compatibility wheel from GitHub Releases, or build Warp from source with CUDA 12 (#1955).
  • Experimental CUDA source builds are available on Windows ARM64. Use CUDA Toolkit 13.4 or newer, an ARM64 Python environment, and the ARM64 Visual Studio C++ tools. CMake builds require CMake 4.4 or newer. Prebuilt wheels and libmathdx are not yet available for this platform. Pass --no-use-libmathdx to build_lib.py. See the source-build guide for details.
  • Windows wheels now include native import libraries. warp.lib, warp-clang.lib, and the new warp_clang.h header let native C++ applications link without resolving each entry point manually. Keep headers, import libraries, and DLLs on the same Warp release (#1888).

Acknowledgments

We also thank the following contributors from outside the core Warp development team:

  • @azrabano23 for fixing scaled sparse expressions and FEM scalar-output validation (#1910, #1912).
  • @elbourne12345 for making signed integer floor division match Python and NumPy (#1918).
  • @HuzaifaAbdulRehman for restoring CPU vertex updates in FEM shape optimization (#1832).
  • @IshaanPotle for correcting documented argument names in FEM and USD rendering APIs.
  • @JaimeFine for clarifying how to run the native APIC visualization examples.
  • @pei-tian for fixing stale node data in FEM partitions and restrictions after rebuilds (#1852).
  • @rickpr for making wp.utils.array_scan() raise a Python exception when a CUDA scan fails (#1894).
  • @TusharND12 for fixing source builds from paths containing spaces on Linux and macOS (#1925).
  • @zhihuidu-amd for fixing Clang CUDA compilation of array slicing and improving HIP graph capture (#1936).

For a complete list of changes, see the full changelog.