You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
num_envs
rays per env
update period
primitive-only scene
mesh scene
hfield scene
no filter vs downstream multi-hop filter
same-process GPU path vs legacy NumPy path
Proposal: backend-neutral ray-cast/LiDAR core with a zero-copy perception fast path
Summary
UniLab needs a perception/ray-cast capability that can support LiDAR-style sensors and InstinctMJ-style depth/ray-camera workloads, but the current Manager-Based observation path is NumPy end-to-end. A direct integration would force every ray result from device memory to host NumPy and then back to a learner GPU. That defeats the main purpose of adding a ray caster for high-throughput training.
This discussion proposes a split that keeps the core generic and puts task-specific semantics downstream:
UniSim owns a backend-neutral ray-query capability and its lifecycle/capability contract.
UniLab owns an optional GPU perception observation fast path that can deliver a backend-owned device buffer to a policy without a D2H/H2D round trip.
Downstream task packages such as instinct_unilab own ray filters, multi-hop logic, camera/LiDAR pattern generation, noise, crop/resize, depth normalization, history, and all task-specific processing.
The important scope boundary is that UniSim/UniLab should not port InstinctMJ's sensor class as-is. They should provide the generic ray-query and device-observation primitives needed to implement that behavior downstream.
1. Feasibility of a backend-neutral UniSim LiDAR contract
A fully backend-neutral contract is feasible if it is defined around ray queries over an explicit collision world, rather than around MuJoCo internals or a particular renderer:
ray origins + directions
|
| backend-owned collision representation
v
raw intersection outputs
The public contract should not expose:
mujoco.MjModel / mujoco.MjData;
mujoco_warp objects;
Warp arrays;
CUDA pointers;
backend-private render contexts;
renderer semantics such as RGB/segmentation.
A suitable neutral output set is:
distance
hit position
normal
geom id / primitive id
body id or collision-shape owner id
hit/miss mask
All outputs should be optional capability features rather than mandatory for every backend. For example, a simple static-map LiDAR backend may only support distance and hit point, while MuJoCo backends can additionally provide geom/body identity.
Required backend profiles
The contract should explicitly distinguish execution profiles rather than pretending all backends are identical.
A. Native backend profile
A backend owns a device-resident collision world and can execute the query directly.
This is the best match for InstinctMJ because InstinctMJ itself is built on mujoco_warp.rays.
B. Pose-sync profile
The backend supplies transforms into a backend-neutral Warp ray runtime. This is the practical route for the current mujoco/mjbatch backend, whose canonical executor state is host-side:
mjbatch host state
-> selected body poses
-> one H2D transform upload
-> immutable collision descriptor + Warp BVH
-> device-resident ray outputs
The transform upload is unavoidable for current mjbatch. The critical requirement is that ray outputs must not be downloaded to host merely to be uploaded again by the learner.
C. External/static-map profile
A backend need not support full physics geometry. It can support a static map or explicitly declared collision proxies plus a small set of moving body poses. This is useful for non-MuJoCo backends and for lightweight LiDAR tasks.
Unsupported geometry, filters, output attributes, or device profiles must fail closed.
What “backend neutral” cannot promise
A universal exact geometry implementation across every physics backend is not realistic. Different backends use different collision representations and may not expose the same primitive/mesh semantics. The neutral contract can promise:
stable input/output semantics;
explicit capability declarations;
deterministic miss/filter metadata behavior;
exact or approximate support levels;
an immutable collision-world descriptor for portable implementations;
no backend-private types crossing the API.
For example:
native MuJoCo can report exact MuJoCo geom ids;
a portable proxy world may report neutral shape ids;
a static map may report no geom/body identity;
deformable meshes and mutable hfields should initially be unsupported unless explicitly represented in the contract.
2. Proposed UniSim core capability
UniSim should add a ray-query capability to SimBackend, similar in spirit to the existing height-scanner capability but designed for arbitrary batched rays and device-resident outputs.
Conceptually:
classBackendRayCaster(Protocol):
deftrace(
self,
origins, # [num_envs, num_rays, 3]directions, # [num_envs, num_rays, 3]*,
max_distance,
env_ids=None,
) ->RayCastFrame:
...
classRayCastFrame:
# Opaque, backend-owned stable frame; never a raw pointer.
...
The exact Python API needs an ADR, but the contract should support:
fixed batch shape at materialization;
per-environment ray origins and directions;
selected-row updates;
update masks/periods owned downstream;
stable output buffers;
explicit output metadata;
explicit host fallback;
explicit device-only profile;
frame versioning and lifetime semantics.
Core query outputs
The core query should return raw physical intersections only:
distance
hit point
normal
geom/body/collision owner id
It should not own:
pinhole camera intrinsics;
depth-image semantics;
LiDAR scan patterns;
Gaussian blur;
crop/resize;
normalization;
temporal history;
latency models;
InstinctMJ mesh-path filters;
multi-hop min-distance filtering;
RGB/segmentation.
Those belong to UniLab perception adapters or downstream task packages.
Collision-world descriptor
For the pose-sync profile, UniSim should materialize an immutable collision descriptor on a cold path:
collision shape kind
local size
local pose
owner body id
mesh id
hfield id / triangulated mesh
include/exclude metadata
per-shape primitive list
The hot path should then synchronize only selected poses:
body position + rotation
The ray kernel can reconstruct world-space collision poses:
This is cheaper and more portable than synchronizing every geom pose. It assumes rigid shapes and immutable local geometry during a step. Reset-time geometry randomization must update the descriptor or declare that configuration unsupported.
3. Zero-copy is a hard requirement, not an optimization flag
A normal NumPy observation path cannot satisfy the goal:
ray kernel on GPU
-> .numpy()
-> ObservationManager
-> learner H2D
That is a D2H/H2D round trip for potentially millions of rays and is not acceptable as the primary perception path.
Device frame contract
UniSim/UniLab should introduce a typed, backend-owned device frame:
__dlpack__ or another explicit zero-copy interchange protocol;
explicit synchronization semantics;
explicit lifetime/retire semantics;
host download only when explicitly requested.
It must not expose:
raw CUDA pointers;
Warp object identity;
a private backend render context;
mutable unversioned buffers.
Zero-copy lifetime design
Aliasing a mutable “current output” buffer is not enough. If the ray caster overwrites it while a policy still reads it, the result is nondeterministic.
The backend should use a small ring or double-buffer and expose frame ownership semantics:
produce frame k
-> downstream acquires view
-> downstream releases/retires view
-> backend may overwrite only after release
For the first implementation, double or triple buffering with conservative stream ordering is preferable to a complex allocator. The contract can later add a more general frame pool.
Same-device requirement
Zero-copy is initially only valid when the ray runtime and learner tensor are on the same process and same CUDA device. The contract must fail closed on device mismatch rather than perform a hidden copy.
Multi-process limitation
Current spawn collectors cannot safely receive a CUDA pointer across a process boundary. Therefore there are two valid phase-1 choices:
run the GPU perception short path in the same process as policy inference; or
explicitly stage through pinned host memory or a designed CUDA IPC mechanism.
Cross-process CUDA IPC is a separate synchronization/lifetime contract and should not be smuggled in as “just pass a pointer.”
If strict zero-copy is required even for existing spawn-based distributed collection, the IPC design must be part of the ADR and implementation scope from the beginning.
4. UniLab GPU perception short path
UniLab needs a perception-specific observation path because the existing ObservationManager, NpEnvState.obs, and uni_rl env contract deliberately assume NumPy arrays. Rewriting every observation term immediately would be too invasive.
The proposal is to add an optional fast path alongside the current path:
downstream task generates rays on GPU
|
backend ray caster executes on same GPU
|
raw device outputs
|
downstream GPU filter / postprocess / history
|
device observation arena
|
policy inference
The device observation arena should be a stable, preallocated batch tensor region. It can contain only perception terms initially; ordinary NumPy observations can continue separately. Mixing paths must be explicit.
Owner split inside UniLab
Manager-Based runtime: optional perception observation lifecycle and row-reset semantics.
Training wrappers: consume a device observation arena and expose it to the selected learner/inference path.
Sim2sim/checkpoint contract: record image/ray observation shape and policy dimensions.
uni_rl: only if its public env contract is generalized to support array-like device observations. Otherwise the first path should be owned by UniLab's direct RL wrapper and avoid making uni_rl import UniLab or a backend SDK.
uni_rl already has DLPack-aware conversion utilities, so a later generic extension is plausible, but its current EnvStateProtocol still declares NumPy arrays. That contract should not be changed casually.
Shape policy
Current obs_groups_spec is flat:
{"obs": int, "critic": int}
For compatibility, the first implementation can expose flattened device perception output:
[num_envs, H * W] or [num_envs, history * H * W]
A native image shape:
[num_envs, channels, H, W]
should be a separate, explicitly designed image-observation contract. It should not be introduced accidentally through one backend-specific sensor.
5. Downstream ownership: instinct_unilab
A downstream package such as instinct_unilab should own all task-specific behavior.
Examples:
pinhole camera intrinsics and pixel-ray generation;
ring/LiDAR scan patterns;
camera/LiDAR frame attachment;
per-env ray-start and ray-direction randomization;
drift and extrinsic randomization;
InstinctMJ-compatible mesh/body selection;
min-distance rejection;
multi-hop ray continuation;
depth-to-image-plane conversion;
clipping semantics;
crop/resize;
Gaussian blur;
depth normalization;
history and latency;
policy observation assembly;
reference-scene parity tests.
UniSim should provide raw ray intersections and sufficient hit metadata. It should not know that a ray set represents a camera or a particular robot's perception model.
This split lets other tasks reuse the same core for:
spinning LiDAR;
depth ray camera;
height-style ray sampling;
obstacle masks;
privileged geometry probes;
static-map navigation sensors.
6. Backend implementation plan
Phase 1: portable mjbatch/Warp runtime
Use the reusable parts of MuJoCo-LiDAR:
core_warp/geometry.py
core_warp/kernels.py
They already contain primitive/mesh intersection and grouped BVH kernels. The existing MjLidarWarp wrapper is not suitable directly because it:
reads MjModel/MjData directly;
allocates Warp arrays per call;
synchronizes per trace;
immediately downloads outputs with .numpy();
assumes a shared ray pattern across envs.
The first integration should turn it into:
immutable MuJoCo collision descriptor
persistent transform buffers
persistent per-env ray buffers
persistent output buffers
grouped BVH
one explicit pose H2D upload per update
device-resident outputs
Required kernel changes include:
per-env ray origins and directions;
optional hit ids/normals;
stable output shape;
no Python-side metadata lookup in step/reset hot paths;
no output readback in the zero-copy profile.
Phase 2: native MjwarpBackend
Implement the same UniSim capability using backend-owned MuJoCo Warp model/data and the ray/render API:
This avoids even the pose H2D upload from the mjbatch profile. It should remain entirely inside the MjwarpBackend; task or training code must not import mujoco_warp private helpers.
CUDA graph integration should wait until buffer shapes and lifecycle are stable. Initial ray operations can execute outside captured physics graphs, followed by a graph-safe version only if it does not introduce hidden synchronization or allocation.
Phase 3: non-MuJoCo backends
Other backends can choose among:
native ray implementation;
pose-sync plus portable collision descriptor;
static external map;
unsupported.
No backend should be allowed to fake support through a silently different collision world.
7. Capability declarations
Rather than a single boolean, support should be declared at semantic granularity. Examples:
num_envs
rays per env
update period
primitive-only scene
mesh scene
hfield scene
no filter vs downstream multi-hop filter
same-process GPU path vs legacy NumPy path
Zero-copy profile acceptance:
no .numpy() call between ray execution and learner consumption;
no unbounded or per-step large allocation;
no implicit host staging on the same-device path;
no hidden transfer on DLPack/device-frame acquisition;
frame overwrite prohibited until release/retire;
device mismatch fails closed.
Legacy profile acceptance:
exactly one explicit D2H per requested host frame;
no repeated D2H through get_sensor_data() / BackendSensorView;
current NumPy env contract remains functional.
9. Recommended work items
ADR
public ray-query capability;
device frame ownership and synchronization;
GPU perception short path;
backend-neutral vs native support levels;
multi-process zero-copy boundary.
UniSim contract PR
ray-caster interfaces;
capability declarations;
SDK-free device-frame protocol;
fake/reference implementation;
contract tests and import-boundary tests.
Portable mjbatch/Warp runtime PR
immutable collision descriptor;
persistent buffers;
per-env rays;
body-pose to collision-pose reconstruction;
device output only;
explicit host fallback.
UniLab perception short-path PR
device observation arena;
downstream GPU term interface;
policy wrapper consumption;
row reset/frame semantics;
legacy fallback.
Downstream instinct_unilab integration
camera ray generation;
filters;
postprocessing;
history;
InstinctMJ parity scenes and tolerances.
10. Decision requested
Do we accept the proposed owner split:
UniSim = generic ray query + capability/device frame;
UniLab = GPU perception short path;
downstream task = ray generation/filter/postprocess?
Is same-process/same-device zero-copy sufficient for the first milestone, with cross-process CUDA IPC designed separately?
Do we start with the mjbatch pose-sync/Warp runtime, then add native MjwarpBackend, or prioritize MjwarpBackend first for InstinctMJ parity?
Do we require flattened observation dimensions first and defer a native image-shape observation contract?
Do we accept strict fail-closed capability declarations for approximate/static-map/proxy implementations?
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
提案:后端中立的 ray-cast/LiDAR 核心与零拷贝感知快路径
概述
UniLab 需要一套感知/ray-cast 能力,用于支持 LiDAR 风格的传感器以及 InstinctMJ 风格的深度/射线相机工作负载。但当前 Manager-Based 的观测路径端到端都是 NumPy。直接集成会迫使每条 ray 结果从设备内存搬到主机 NumPy,再搬回 learner GPU。这就违背了为高吞吐训练引入 ray caster 的主要目的。
本讨论提出一种拆分方式:核心保持通用,任务特定语义下沉到下游:
instinct_unilab)拥有 ray 过滤器、multi-hop 逻辑、camera/LiDAR 模式生成、噪声、crop/resize、深度归一化、history 以及所有任务特定处理。重要的范围边界是:UniSim/UniLab 不应照搬 InstinctMJ 的 sensor 类,而应提供在下游实现该行为所需的通用 ray-query 与设备观测原语。
1. 后端中立的 UniSim LiDAR 契约的可行性
如果契约围绕显式碰撞世界上的 ray query 来定义,而不是围绕 MuJoCo 内部结构或某个特定渲染器,那么完全后端中立的契约是可行的:
公开契约不应暴露:
mujoco.MjModel/mujoco.MjData;mujoco_warp对象;合适的中立输出集合是:
所有输出都应是可选的 capability 特性,而非每个后端都必须支持。例如,一个简单的静态地图 LiDAR 后端可能只支持 distance 和 hit point,而 MuJoCo 后端可以额外提供 geom/body 标识。
必需的后端 profile
契约应明确区分不同的执行 profile,而不是假装所有后端都一样。
A. 原生后端 profile
后端拥有驻留设备的碰撞世界,可直接执行查询。
最终期望的实现:
这与 InstinctMJ 最匹配,因为 InstinctMJ 本身就构建在
mujoco_warp.rays之上。B. Pose-sync profile
后端将变换(transform)提供给一个后端中立的 Warp ray 运行时。这是当前
mujoco/mjbatch后端的实际路线,因为其规范的执行器状态在主机侧:对当前的 mjbatch 来说,transform 上传不可避免。关键要求是:ray 输出绝不能仅仅为了被 learner 再次上传而先下载到主机。
C. 外部/静态地图 profile
后端不必支持完整的物理几何。它可以支持静态地图,或显式声明的碰撞代理(collision proxy)加少量运动 body pose。这对非 MuJoCo 后端和轻量级 LiDAR 任务很有用。
不支持的几何、过滤器、输出属性或设备 profile 必须 fail closed(拒绝而非静默降级)。
“后端中立”无法承诺的东西
跨所有物理后端实现统一的精确几何并不现实。不同后端使用不同的碰撞表示,可能不暴露相同的 primitive/mesh 语义。中立契约可以承诺的是:
例如:
2. 拟议的 UniSim 核心能力
UniSim 应在
SimBackend上增加 ray-query 能力,理念上类似现有的 height-scanner 能力,但面向任意批量 ray 和设备驻留输出。概念上:
具体的 Python API 需要一份 ADR,但契约应支持:
核心查询输出
核心查询应只返回原始物理求交结果:
它不应负责:
这些属于 UniLab 感知适配器或下游任务包。
碰撞世界描述符
对于 pose-sync profile,UniSim 应在冷路径上 materialize 一个不可变的碰撞描述符:
热路径随后只同步选定的 pose:
ray kernel 可以重建世界系下的碰撞 pose:
这比同步每个 geom pose 更便宜也更可移植。它假设形状为刚体,且局部几何在一个 step 内不可变。reset 时的几何随机化必须更新描述符,或声明该配置不受支持。
3. 零拷贝是硬性要求,不是优化开关
普通的 NumPy 观测路径无法满足目标:
对潜在数百万条 ray 来说,这是一次 D2H/H2D 往返,作为主感知路径是不可接受的。
设备帧契约
UniSim/UniLab 应引入一个带类型的、后端持有的设备帧:
它应暴露:
__dlpack__或其他显式的零拷贝交换协议;它不得暴露:
零拷贝生命周期设计
仅仅别名化一个可变的“当前输出”缓冲是不够的。如果 ray caster 在 policy 仍在读取时覆写它,结果将是不确定的。
后端应使用小型环形缓冲或双缓冲,并暴露帧所有权语义:
第一个实现应优先采用双缓冲或三缓冲加保守的流序保证,而不是复杂的分配器。契约之后可以再增加更通用的 frame pool。
同设备要求
零拷贝初期仅在 ray 运行时与 learner tensor 位于同一进程、同一 CUDA 设备时有效。契约必须在设备不匹配时 fail closed,而不是执行隐式拷贝。
多进程限制
当前的 spawn collector 无法安全地跨进程边界接收 CUDA 指针。因此第一阶段有两个合法选择:
跨进程 CUDA IPC 是另一套同步/生命周期契约,不应以“只是传个指针”的方式偷偷混入。
如果现有的基于 spawn 的分布式采集也严格要求零拷贝,那么 IPC 设计必须从一开始就属于 ADR 与实现范围的一部分。
4. UniLab GPU 感知短路径
UniLab 需要一条感知专用的观测路径,因为现有的
ObservationManager、NpEnvState.obs和uni_rlenv 契约都有意假定使用 NumPy 数组。立刻重写所有观测 term 侵入性太大。提案是在现有路径之外增加一条可选快路径:
拟议流水线
设备观测 arena 应是一块稳定的、预分配的 batch tensor 区域。初期它可以只包含感知 term;普通 NumPy 观测可继续走原路径。两条路径的混用必须是显式的。
UniLab 内部的所有权划分
uni_rlimport UniLab 或后端 SDK。uni_rl已有 DLPack 感知的转换工具,因此之后做通用扩展是可行的,但其当前的EnvStateProtocol仍声明 NumPy 数组。该契约不应被轻易改动。Shape 策略
当前的
obs_groups_spec是扁平的:{"obs": int, "critic": int}为兼容性考虑,第一个实现可以暴露扁平化的设备感知输出:
原生图像 shape:
应是一份单独的、显式设计的图像观测契约。不应通过某个后端特定 sensor 意外引入。
5. 下游所有权:
instinct_unilab诸如
instinct_unilab的下游包应拥有所有任务特定行为。例如:
UniSim 应提供原始 ray 求交结果和足够的命中元数据。它不应知道一组 ray 代表一台相机或某个机器人的感知模型。
这种拆分让其他任务可以复用同一核心来实现:
6. 后端实现计划
阶段 1:可移植的 mjbatch/Warp 运行时
复用 MuJoCo-LiDAR 中可复用的部分:
它们已包含 primitive/mesh 求交和分组 BVH kernel。现有的
MjLidarWarp封装不适合直接使用,因为它:MjModel/MjData;.numpy()下载输出;第一次集成应将其改造为:
所需的 kernel 修改包括:
阶段 2:原生
MjwarpBackend使用后端持有的 MuJoCo Warp model/data 和 ray/render API 实现同一 UniSim 能力:
这连 mjbatch profile 中的 pose H2D 上传都可以省掉。它应完全留在
MjwarpBackend内部;task 或 training 代码不得 importmujoco_warp的私有 helper。CUDA graph 集成应等到缓冲 shape 与生命周期稳定之后。初期的 ray 操作可以在被捕获的物理 graph 之外执行,仅当不会引入隐藏同步或分配时,才跟进 graph-safe 版本。
阶段 3:非 MuJoCo 后端
其他后端可以在以下方案中选择:
不允许任何后端通过静默使用不同碰撞世界的方式伪装支持。
7. Capability 声明
支持度不应是单一布尔值,而应按语义粒度声明。例如:
当下游任务要求不受支持的特性时,配置应 fail closed。
Capability 报告应区分:
例如,代理碰撞世界不应声称具有精确的 MuJoCo 几何标识。
8. 性能验收标准
实现与评审应测量完整路径,而不仅是求交 kernel。
最小计时分段:
必需的工作负载维度:
零拷贝 profile 验收标准:
.numpy()调用;Legacy profile 验收标准:
get_sensor_data()/BackendSensorView重复 D2H;9. 建议的工作项
ADR
UniSim 契约 PR
可移植 mjbatch/Warp 运行时 PR
UniLab 感知短路径 PR
下游
instinct_unilab集成10. 请求决策
MjwarpBackend,还是为了 InstinctMJ 一致性优先做MjwarpBackend?English original(英文原文)
Proposal: backend-neutral ray-cast/LiDAR core with a zero-copy perception fast path
Summary
UniLab needs a perception/ray-cast capability that can support LiDAR-style sensors and InstinctMJ-style depth/ray-camera workloads, but the current Manager-Based observation path is NumPy end-to-end. A direct integration would force every ray result from device memory to host NumPy and then back to a learner GPU. That defeats the main purpose of adding a ray caster for high-throughput training.
This discussion proposes a split that keeps the core generic and puts task-specific semantics downstream:
instinct_unilabown ray filters, multi-hop logic, camera/LiDAR pattern generation, noise, crop/resize, depth normalization, history, and all task-specific processing.The important scope boundary is that UniSim/UniLab should not port InstinctMJ's sensor class as-is. They should provide the generic ray-query and device-observation primitives needed to implement that behavior downstream.
1. Feasibility of a backend-neutral UniSim LiDAR contract
A fully backend-neutral contract is feasible if it is defined around ray queries over an explicit collision world, rather than around MuJoCo internals or a particular renderer:
The public contract should not expose:
mujoco.MjModel/mujoco.MjData;mujoco_warpobjects;A suitable neutral output set is:
All outputs should be optional capability features rather than mandatory for every backend. For example, a simple static-map LiDAR backend may only support distance and hit point, while MuJoCo backends can additionally provide geom/body identity.
Required backend profiles
The contract should explicitly distinguish execution profiles rather than pretending all backends are identical.
A. Native backend profile
A backend owns a device-resident collision world and can execute the query directly.
Expected eventual implementation:
This is the best match for InstinctMJ because InstinctMJ itself is built on
mujoco_warp.rays.B. Pose-sync profile
The backend supplies transforms into a backend-neutral Warp ray runtime. This is the practical route for the current
mujoco/mjbatchbackend, whose canonical executor state is host-side:The transform upload is unavoidable for current mjbatch. The critical requirement is that ray outputs must not be downloaded to host merely to be uploaded again by the learner.
C. External/static-map profile
A backend need not support full physics geometry. It can support a static map or explicitly declared collision proxies plus a small set of moving body poses. This is useful for non-MuJoCo backends and for lightweight LiDAR tasks.
Unsupported geometry, filters, output attributes, or device profiles must fail closed.
What “backend neutral” cannot promise
A universal exact geometry implementation across every physics backend is not realistic. Different backends use different collision representations and may not expose the same primitive/mesh semantics. The neutral contract can promise:
For example:
2. Proposed UniSim core capability
UniSim should add a ray-query capability to
SimBackend, similar in spirit to the existing height-scanner capability but designed for arbitrary batched rays and device-resident outputs.Conceptually:
The exact Python API needs an ADR, but the contract should support:
Core query outputs
The core query should return raw physical intersections only:
It should not own:
Those belong to UniLab perception adapters or downstream task packages.
Collision-world descriptor
For the pose-sync profile, UniSim should materialize an immutable collision descriptor on a cold path:
The hot path should then synchronize only selected poses:
The ray kernel can reconstruct world-space collision poses:
This is cheaper and more portable than synchronizing every geom pose. It assumes rigid shapes and immutable local geometry during a step. Reset-time geometry randomization must update the descriptor or declare that configuration unsupported.
3. Zero-copy is a hard requirement, not an optimization flag
A normal NumPy observation path cannot satisfy the goal:
That is a D2H/H2D round trip for potentially millions of rays and is not acceptable as the primary perception path.
Device frame contract
UniSim/UniLab should introduce a typed, backend-owned device frame:
It should expose:
__dlpack__or another explicit zero-copy interchange protocol;It must not expose:
Zero-copy lifetime design
Aliasing a mutable “current output” buffer is not enough. If the ray caster overwrites it while a policy still reads it, the result is nondeterministic.
The backend should use a small ring or double-buffer and expose frame ownership semantics:
For the first implementation, double or triple buffering with conservative stream ordering is preferable to a complex allocator. The contract can later add a more general frame pool.
Same-device requirement
Zero-copy is initially only valid when the ray runtime and learner tensor are on the same process and same CUDA device. The contract must fail closed on device mismatch rather than perform a hidden copy.
Multi-process limitation
Current spawn collectors cannot safely receive a CUDA pointer across a process boundary. Therefore there are two valid phase-1 choices:
Cross-process CUDA IPC is a separate synchronization/lifetime contract and should not be smuggled in as “just pass a pointer.”
If strict zero-copy is required even for existing spawn-based distributed collection, the IPC design must be part of the ADR and implementation scope from the beginning.
4. UniLab GPU perception short path
UniLab needs a perception-specific observation path because the existing
ObservationManager,NpEnvState.obs, anduni_rlenv contract deliberately assume NumPy arrays. Rewriting every observation term immediately would be too invasive.The proposal is to add an optional fast path alongside the current path:
Proposed pipeline
The device observation arena should be a stable, preallocated batch tensor region. It can contain only perception terms initially; ordinary NumPy observations can continue separately. Mixing paths must be explicit.
Owner split inside UniLab
uni_rlimport UniLab or a backend SDK.uni_rlalready has DLPack-aware conversion utilities, so a later generic extension is plausible, but its currentEnvStateProtocolstill declares NumPy arrays. That contract should not be changed casually.Shape policy
Current
obs_groups_specis flat:{"obs": int, "critic": int}For compatibility, the first implementation can expose flattened device perception output:
A native image shape:
should be a separate, explicitly designed image-observation contract. It should not be introduced accidentally through one backend-specific sensor.
5. Downstream ownership:
instinct_unilabA downstream package such as
instinct_unilabshould own all task-specific behavior.Examples:
UniSim should provide raw ray intersections and sufficient hit metadata. It should not know that a ray set represents a camera or a particular robot's perception model.
This split lets other tasks reuse the same core for:
6. Backend implementation plan
Phase 1: portable mjbatch/Warp runtime
Use the reusable parts of MuJoCo-LiDAR:
They already contain primitive/mesh intersection and grouped BVH kernels. The existing
MjLidarWarpwrapper is not suitable directly because it:MjModel/MjDatadirectly;.numpy();The first integration should turn it into:
Required kernel changes include:
Phase 2: native
MjwarpBackendImplement the same UniSim capability using backend-owned MuJoCo Warp model/data and the ray/render API:
This avoids even the pose H2D upload from the mjbatch profile. It should remain entirely inside the
MjwarpBackend; task or training code must not importmujoco_warpprivate helpers.CUDA graph integration should wait until buffer shapes and lifecycle are stable. Initial ray operations can execute outside captured physics graphs, followed by a graph-safe version only if it does not introduce hidden synchronization or allocation.
Phase 3: non-MuJoCo backends
Other backends can choose among:
No backend should be allowed to fake support through a silently different collision world.
7. Capability declarations
Rather than a single boolean, support should be declared at semantic granularity. Examples:
Configuration should fail closed when a downstream task requires an unsupported feature.
The capability report should distinguish:
For example, a proxy collision world should not claim exact MuJoCo geometry identity.
8. Performance acceptance criteria
The implementation and review should measure the full path, not only the intersection kernel.
Minimum timing segments:
Required workload axes:
Zero-copy profile acceptance:
.numpy()call between ray execution and learner consumption;Legacy profile acceptance:
get_sensor_data()/BackendSensorView;9. Recommended work items
ADR
UniSim contract PR
Portable mjbatch/Warp runtime PR
UniLab perception short-path PR
Downstream
instinct_unilabintegration10. Decision requested
MjwarpBackend, or prioritizeMjwarpBackendfirst for InstinctMJ parity?All reactions