Releases: LancerLSY/sentinel-evc-lab
Release list
Sentinel Launch Gate: VLA integration checks before motion
Sentinel Launch Gate checks VLA integration changes before the first command. Fixed inputs pass through the actual adapter; the product compares named camera tensors, ordered state values, the selected action cursor and exact final action bytes. It blocks changed behavior at the trusted software writer and shows the first mismatch plus observed camera, component and conversion clues.
CLI and the native macOS App export a signed decision capsule. A CPU machine can authenticate and recompute the recorded decision without model weights. The project page lets visitors inspect each retained input and fault case.
Measured behavior
- RTX 4090 D (24 GB): 40 fresh official SmolVLA inference paths over ten LIBERO Spatial initial inputs. All 6/6 injected integration configurations block before writer entry; blocked downstream calls are zero. The unchanged integration passes. No environment stepping occurs in this launch experiment.
- Signed-record adapter replay: 10/10 injected integration configurations block on 20 requests. Identity-only behavior-preserving change passes; incomparable input returns REVIEW.
- All 13 recorded and 7 GPU capsules reproduce. The GPU camera-swap capsule is 26,172 bytes (25.6 KiB) before compression; its independent public key is published separately.
- The actual native App imported the GPU probe files, displayed the camera-route BLOCK, saved ZIP and public key through native dialogs, and reproduced the saved decision in CLI.
- Existing regression: 211 passed, 0 failed, 9 skipped. Exact GPU runner and qualifier source hashes match the published code.
The installed KineGrant 2.65.5 and RLSOK 1.5.12 comparison retains all correctly authorized calls and RLSOK release-difference results. Sentinel's addition is the packaged observed-behavior launch check, specific mismatch/repair clues and signed decision reproduction. Ordinary authorization remains tied; these results do not claim physical safety or global exclusivity.
Use and reproduce
Project page · CLI/App integration · Full outcomes and protocols
Download the evidence archive and artifact index, check SHA-256, and follow its README. The source archive freezes the exact public commit. The complete page archive includes retained 3D models, replays, dual-camera videos and the new launch-check explorer.
PASS covers the retained finite inputs. Intentional behavior changes need a newly reviewed reference. Model inference, task success and physical hardware require their respective validation. The local App remains a locally built developer preview.
Verified execution differences: App, CLI and paired 3D evidence
Verified execution differences are now available in the local CLI and macOS App. Compare independently keyed recordings, identify the first changed VLA field, open both source frames and apply an explicit regression gate.
Across the complete 40-episode pair, all 5,420 actions agree; the product finds one camera-2 feedback difference at task 0/state 48/step 4. The installed Rerun 0.38.1 CLI detects the differing episode chunk; Sentinel additionally supplies the first step, field, expected authorization classification and source-frame navigation. Raw commands and outputs are included. No cross-tool speed ranking is claimed.
The project page includes paired original Panda motion, a synchronized recorded-step slider, all 40 diagnostics and explicit labels for missing step-4 body poses.
Assets: frozen source, complete portable project page, diagnosis/RRD evidence and a size/SHA-256 index. The original signed recordings and independent public keys remain in the native product release. This remains a local developer preview; physical robot diagnostics are read-only.
Native writer entry: exact arrays and installed authorization comparisons
The native gateway now preserves an owned action copy and rechecks live feedback, context and timing at writer entry. When a caller changes its array during preparation, the approved bytes still reach dispatch. In the fixed concurrency grid, invalid downstream dispatches fall from 40 to 0; all ten caller-buffer cases execute the approved copy.
Installed KineGrant 2.65.5 and RLSOK 1.5.12 each pass the same 60/60 ordinary authorization fault cases as Sentinel. Real CycloneDDS Security delivers 70/70 frozen records. The evidence pack retains individual outcomes, controls, adapters, qualification output, signatures and corrections. Exploratory callbacks are reported separately from the common-input grid.
These recorded-action software experiments establish the integration contract. They do not establish global exclusivity, physical robot safety or an updated gateway latency ranking.
Verified motion reuse: UR5e timing, write permissions and 3D evidence
Sentinel reuses valid motion checks after an action chunk changes and binds approved bytes to the software write boundary.
The frozen four-candidate UR5e experiment measured 1.092–1.619× speedup over full continuous checking, including root and issued-certificate cost. The ten-cell grid retained 249,120 unique candidates with zero incremental/full decision differences. Single-candidate checking was slower, and the largest batches retained 50 ms deadline misses.
The numeric CLI/App lifecycle now reuses exact remaining-suffix certificates. The UR5e geometry experiment is separate from native VLA execution. Timing used Xeon CPU; the experiment host GPU was RTX 4090 D 24 GB.
Assets include the complete frozen candidate traces, raw timings, protocols and root-cluster analysis, fault drills, independent MuJoCo contact checks, signed numeric product records, reviewed source, and a 1080p collision replay. All public asset hashes are listed in artifact-index.json.
Simulator replay evidence
Simulator replay evidence
Exact-action diagnostic replay of LIBERO-Spatial task 5, initial state 21, seed 43022. MuJoCo 3.8.1 reproduces failure after 280 actions (reward 0); MuJoCo 3.3.7 reproduces success after 85 actions (reward 1). Initial observations, dispatched actions and official outcomes match each frozen branch. No policy inference or observer intervention; excluded from benchmark denominators.
The 1080p bilingual film has 280 frames at 20 fps (14 seconds). The successful pane holds its terminal frame for 195 frames, explicitly labelled as display only. Full H.264 decode passes.
Source, raw trajectories, first camera images, capture manifests, runtime receipt and per-file artifact hashes are retained. Official simulator asset content is independently verified at a pinned revision: 239 files, 189,311,115 bytes. This portable content binding does not claim reconstruction of the historical cache-containing directory checksum. Capture source freeze: f59f81fda07665f873076f4a4e8e31918d2f7120. Hardware: RTX 4090 D (24 GB).
The original paper-validation evidence, including all failed benchmark episodes and four earlier films, remains separately retained. This comparison supports simulator-version compatibility differences; it does not isolate a causal mechanism or qualify real-robot performance.
Sentinel EVC · paper validation, native rollouts and 3D evidence
Sentinel EVC: native rollouts, physical stress and final-action verification
This research release combines frozen experimental designs, complete raw outcomes, reproducible source, result-derived figures and four actual simulator films. All measured GPU work uses RTX 4090 D (24 GB).
Full report · 中文报告 · Experimental matrix · Mechanism review
| Study | Retained result | Interpretation |
|---|---|---|
| Full-suite native confirmation, 100 fresh-state rollouts | 93/100 success, descriptive Wilson 95% interval 86.25–96.57%, zero crashes | Official SmolVLA/Panda checkpoint, isolated MuJoCo 3.3.7, execute 10 of 50 predicted actions |
| Complete backend comparison, 40 rollouts / 20 matched pairs | Task 5: 0/10 versus 7/10; control 10/10 versus 10/10 | MuJoCo 3.8.1 versus 3.3.7; benchmark compatibility effect, not an isolated reset cause |
| Original native configuration, 100 rollouts | 58/100 success, all 42 failures retained | Separate states/backend/horizon; the difference from 93/100 is not a paired improvement estimate |
| Execution-horizon comparison, 40 rollouts | Task 5: 0/10 → 1/10, paired interval includes zero | More frequent re-observation alone leaves the severe failure largely unresolved |
| No-policy reset audit, 40 resets | Task 5 target-bowl position differs 35.327 mm across paired versions | Version-sensitive conditions before the first policy action |
| Physical stress, 324 roots / 27 cells | 4.8 s: 324/324 complete, zero observed drops; 1.6 s: 184/324 complete, 78 drops | Slower physical comparator takes 3× the motion duration; no adaptive advantage over fixed 4.8 s |
| Support routing, 600 roots | 300 complete, 300 reject, zero observed unsafe selections | Fixed 4.8 s completes all 600; rejection remains incomplete |
| Hidden future friction, 100 pairs | Fast-action risk differs in 100/100 identical-input pairs | An unobserved future contact change requires a trusted bound or additional sensing |
| UR5e final validation, 60 roots | Full/incremental agree 60/60; mean marginal saving 0.996 ms | Small mean saving; incremental P95 is worse and hazard diversity is limited |
The fresh-state confirmation keeps all seven 280-step failures: four in the top-drawer task, one each in tasks 3, 5 and 8. Top-drawer success remains 6/10. All 11,592 official step outcomes and environment actions are retained, with zero postprocessor-to-environment action difference and zero observer interventions. Its source and protocol were frozen before formal execution at 9b1a48d.
Four simulator films
- sentinel-ur5e-paper.mp4 — full official UR5e meshes, 1920×1080, 19.95 s. Explains safe execution, a rejected suffix change and changed context using retained MuJoCo trajectories. The red proposed motion is counterfactual and was not dispatched.
- friction-failure-vs-fallback.mp4 — 1920×1080, 10.2 s. The same low-friction initial state produces a payload drop with 1.6 s motion and completion with 4.8 s motion; captions and displacement curves show the measured threshold.
- libero-task00-state00-explained.mp4 — predeclared original task 0/state 0, 83 actions at the actual 20 Hz clock, 4.15 s. Native 360×360 RGB is enlarged in a captioned 1080p canvas. It belongs to the original 58/100 configuration.
- task5_state0_failure_bilingual_20hz.mp4 — exact replay of all 280 original failed actions, 14.05 s including the initial frame. Zero reward and no success; target-bowl center rises at most 1.674 mm. Object-pose records and raw RGB are retained. This post-hoc diagnostic adds no benchmark episode.
The optional paired-backend film renderer is included as statically reviewed source only; no executed paired film is claimed or included.
Download and reproduce
- sentinel-paper-validation-data.tar.gz contains every formal root/rollout, paired outcomes, action traces, poses, exact consumed preflight receipts, media, and provenance.
contents_manifest.jsonindexes each retained file by size and SHA-256. - sentinel-paper-validation-source.zip is the exact source of this release tag, including pinned protocols, runners, reproduction commands, bilingual reports and figures.
- artifact_index.json records the source commit and sizes/SHA-256 values of the data, source and individual films.
Native VLA reproduction · Routing and physical reproduction
The official lerobot/smolvla_libero checkpoint evaluated here is distinct from the published Sentinel SO100 overlay. These experiments do not establish that overlay's robot success, Sentinel intervention benefit on native VLA, continuous-domain safety, or physical hardware performance. Existing trained WorldGuard artifacts remain in the earlier GPU/simulation releases. Independent study denominators are never pooled.
Portable product page and CDN-compatible native replay
The project replay loader verifies both the retained compressed-model identity and the decoded source-model SHA-256, with bounded loading. All four task replays require model and action-record identities. This supports CDN transport compression while preserving the original task models and experimental records.
sentinel-project-page-v2.zip contains the complete bilingual project page, four interactive native trajectories, task-specific meshes, dual-camera videos and upstream notices. sentinel-native-vla-source-v2.zip contains the public source snapshot at 834e77e. release-asset-index-v2.json binds both current downloads to their source commit, byte sizes and SHA-256. Earlier delivery assets remain available as prior versions.
Live project page · Original signed runs, data and results · Measured product behavior
The formal 40-pair outcomes, six fault classes, exact action records and audited videos are unchanged. This delivery adds no experiment samples or model training. Download the page archive and extract at the repository root to run a complete local preview of website/.
Native VLA gateway: paired rollouts and recorded 3D evidence
Sentinel authorizes the exact final VLA request immediately before the environment writer and retains a portable signed action/feedback record. CLI and macOS App share the local workbench; the App imports, inspects and exports native runs with an independently selected public key.
- Transparent integration: 40/40 paired action sequences are byte-identical; direct baseline and active authorization both achieve 33/40 task successes over 5,420 exact requests.
- Invalid authorization: all 60/60 synthetic attempts are rejected before the instrumented writer, with zero forbidden writer calls.
- Measured gateway cost: authorization plus admission averages 0.802 ms, P95 1.040 ms, excluding policy inference and environment stepping.
- Inspectable motion: three predeclared successful Panda episodes and the first ordered failure have task meshes, audited dual-camera videos, clear request/permit/outcome captions and source identities.
Hardware: RTX 4090 D (24 GB). The fixed engineering grid uses actual official SmolVLA/LIBERO inference in both lanes. All seven task failures and earlier compatibility/serialization failures are retained. This native profile provides in-process request integrity; physical robot motion and Panda collision/braking validation remain separate work. The App is a locally built developer preview.
Project page · Full results and frozen design · Media and failure interpretation · Published SO100 overlay
Downloads: native-vla-results-pack.zip contains the original signed runs, separate public-key files, reports, frozen configurations, negative results, audited pose/model data, videos and notices. sentinel-native-vla-source.zip is the exact public source snapshot. sentinel-project-page.zip is the complete portable project page, including its large replay assets. Individual native-v3-*.zip and native-v3-*.public files support separate import and verification. release-asset-index.json records every asset's size and SHA-256.
The native Panda benchmark uses the official lerobot/smolvla_libero checkpoint. The published SO100 overlay is a separate recorded-action model. Signature verification establishes integrity relative to the independently selected key; it does not establish sensor truth or a trusted hardware authority. Companion rendering adds no benchmark samples. Upstream mesh/data notices are preserved.
Simulation research: UR5e validation, WorldGuard stress and reproducible data
Transformed robot plans need a decision on the final executable plan. This research release includes bound static-prefix reuse with full fallback, and a separately implemented higher-resolution MuJoCo replay of all 180 constructed UR5e cases.
Measured results:
- UR5e: 139 unsafe / 41 safe constructed roots. Parent-only allowed 125 unsafe roots; full and repaired incremental gates allowed none. 106 static-prefix reuses and 74 full fallbacks; dynamic rollout remains complete.
- WorldGuard: 600 fresh roots across six scenarios. Both learned gates fail on all 100 low-friction roots; the visual gate also allows 7/100 unsafe camera-shift roots. These negative results are retained.
- Separate known-profile fallback: a 4.8 s candidate has 0/100 unsafe roots and 100/100 tray endpoint tasks, at 3x the 1.6 s motion duration. The friction floor is assumed, not sensed; this is not a deployed neural-model repair.
- SmolVLA: previously evaluated five-episode normalized action MAE improves by 63.46%. The model overlay, loader and model card are available in sentinel-smolvla-hub-ready.tar.gz; a development example verifies bitwise-equal loading.
Downloads include the exact source ZIP, retained old/new experiment data and results, the complete pinned Apache-2.0 SO100 dataset, a portable BSD-3-Clause UR5e model, a 33.75 s 720p montage, and the load-verified SmolVLA overlay. Each archive contains individual file hashes; artifact_index.json records asset hashes.
Scope: constructed MuJoCo fixtures and offline recorded-data evaluation. Timing is fixed-order diagnostic with early exits, with no incremental speedup claim. Original UTC freeze-label errors and old evaluator limitations are preserved with errata. No external preregistration, continuous collision certificate, general safety accuracy or physical robot execution is claimed. Measured GPU: RTX 4090 D (24 GB).
GPU experiments 2026-10-02 — WorldGuard and SmolVLA
WorldGuard ensemble comparisons and SmolVLA fine-tuning on SO100 PickPlace recordings, with held-out metrics, pinned inputs and reproducible model artifacts. GPU measurements use an RTX 4090 D (24 GB).
- WorldGuard: 36 trained numerical, MuJoCo object and real SO100 state/visual ensemble members, with calibration, normalization/PCA, predictions, logs and exact source identities.
- SmolVLA: 5000 updates on the pinned real SO100 dataset; dev-selected checkpoint at step 3750. Held-out normalized action MAE decreased from 0.673166 to 0.245984 (63.46%) across five held-out episodes. This is action reconstruction, not hardware task success.
- Batch-one, four-thread image-in-memory to 50x6 action chunk inference: P50/P95 229.26/236.90 ms, excluding video decoding, transport and actuation.
- Camera-layout stress: the original profile produced 47/500 unsafe selections. Profile-matched dev/cal recalibration changed the same predictions to slower 1.6-second selections and observed 0/500 on that existing stress set; XY coverage remained 94.0%. This recovery is not a new independent deployment test.
- Negative comparisons, strong physical/state baselines, actual units and independent sample denominators are retained in the report.
Download artifact_index.json to verify the two bundle SHA-256 digests. Each bundle includes its own per-file ARTIFACT_INDEX.json, loading/model cards, attribution and licenses. The SmolVLA bundle is the selected trainable overlay and requires the pinned official frozen base/backbone; optimizer/RNG and full upstream weights are excluded. Original real videos remain at their dataset source.
Report: https://github.com/LancerLSY/sentinel-evc-lab/blob/main/docs/gpu_training_results.md
Model cards and loading: https://github.com/LancerLSY/sentinel-evc-lab/blob/main/docs/gpu_model_cards.md
Experimental profile. Live robot actuation, real object supervision and product permit/runtime integration remain separate work. Custom code/numerical-WorldGuard weights: MIT; upstream SmolVLA/SmolVLM2 and dataset: Apache-2.0, included with attribution.