SENTINEL EVC

ROBOT POLICY VERIFICATION · OPEN RESEARCH

Sentinel EVCVerify the final robot action.

An execution gate for transformed robot plans, with single-use permits and independently verifiable evidence.

Sentinel EVC Lab · LancerLSY · Personal homepage · Research overview

01 / DEMONSTRATION

The plan changed. The decision must change.

UR5e full-mesh MuJoCo replay · 1080p · 19.95 s. Three labelled illustrations: safe control, changed suffix and changed scene. Rejected motion is explicitly counterfactual and not dispatched.Media provenance
Same initial state, two durations · 1080p · 10.2 s. The 1.6 s motion slips and drops the payload; 4.8 s completes the transport at 3× duration. The graph shows payload displacement and the risk threshold. Actual stored MuJoCo contact states.Media provenance
Official SmolVLA · native LIBERO/Panda · predeclared task 0/state 0 succeeds in 83 actions. Original 360×360 RGB in a 1080p caption canvas, played at the actual 20 Hz clock. Original MuJoCo 3.8.1 / execute-50 study: 58/100; observer makes no interventions. The fresh compatibility configuration is reported separately below.Media provenance
Original task 5/state 0 failure · exact 280-action replay · 720p caption canvas · 14.05 s. Initial state/cameras and every action match the original trace. Target bowl center rises at most 1.674 mm; reward stays zero. Post-hoc diagnostic, excluded from benchmark counts. Distance labels use body centers and observed EEF position.Media provenance
View failure diagnostic: backend comparison
Task 5 · initial-state index 21 · seed 43022 · exact formal-action replay. MuJoCo 3.8.1 fails after 280 steps with reward 0; MuJoCo 3.3.7 succeeds after 85 steps with reward 1. Each backend retains its own formal camera hash. The successful side then holds its terminal display for 195 frames (9.75 s) only to synchronize the film; the hold adds no action, physics step, reward or benchmark sample.Media provenance

02 / OVERVIEW

A check belongs at the point of execution.

Robot policies produce action blocks that may be retimed, repaired, transformed between frames or spliced before execution. A verdict on the original plan does not automatically cover the resulting motion. Sentinel EVC binds the verification decision to the final action and current execution context.

The local research workbench combines final-plan checks, bounded certificate reuse with full fallback, a single writer to the controller and signed evidence. Separate research profiles evaluate UR5e motion, WorldGuard consequence prediction and SmolVLA action reconstruction on recorded SO100 data.

Four separate model, embodiment and evaluation paths
324/324slow-motion completion / stratified contact roots
93/100official SmolVLA success / fresh-state compatibility configuration
600fresh MuJoCo roots / six distribution-shift scenarios

03 / METHOD

Bind the decision to what actually executes.

The action is the binding.

Plans, model/profile identity and current context travel together. A modified suffix, changed scene or stale observation cannot inherit an unrelated verdict.

Reuse has a boundary.

The UR5e study reuses only a matching static prefix. Dynamic rollout always starts from frame zero; an invalid parent record triggers full validation.

Execution leaves evidence.

Submitted, accepted and observed commands are distinct. Revocation blocks new old-generation submissions; recorded events and digests support independent verification.

The product core supports bounded numeric and fixed contact profiles. Trained neural models and the UR5e study are separate research profiles; a live robot motion adapter is not enabled.

Read the mechanism and design trace

04 / RECORDED 3D REPLAY

Inspect the motion behind the verdict.

A schematic replay of the recorded UR5e joint trajectory. The full rendered demonstration is available above.
Drag to orbit · scroll to zoom
0.05 / 2.05 s

Actual saved joint positions, forwarded through the pinned MuJoCo model. Schematic link geometry, without CAD meshes or new dynamics. These two examples were selected after review; outcome labels apply to the whole trajectory, not the current frame.

05 / EXPERIMENTS

Compare the decision, the error and the cost.

Select a research figure to open its full-size version. Counts and evidence scope are explained below each figure.

Prospective support routing: safety and task utility

600 new roots · 100 paired counterfactuals

The protocol and source hashes were committed before formal runs. Under a declared friction floor, the long action completes 100/100 low-friction tasks; camera mismatch routes to the camera-independent state model. Unsupported mass and goal profiles still reject 300/600 roots.

All-root completion, unsafe execution and rejection counts for six scenes and three policies

The router completes 300/600; fixed 4.8 s completes 600/600. Rejection is incomplete. Zero unsafe observations per 100-root scene has a 3.70% Wilson upper bound. Trusted profile declarations do not measure friction.

324-root physical stress: slow fallback completion and fast failure by friction bin

A separately frozen 27-cell stress study evaluates 324 new roots. Fixed 4.8 s completes 324/324 with no observed unsafe/drop events; 1.6 s completes 184/324 and drops 78. The floor rule equals fixed 4.8 s on every root, at 3× duration. Descriptive 95% Wilson zero-event upper bounds: aggregate 1.17%; each cell n=12: 24.25%. This does not qualify the continuous parameter domain.

Small marginal UR5e cost difference and worse incremental P95

On 60 predefined UR5e roots, with obstacles independently sampled before plans, full and incremental decisions agree. Mean marginal saving is 0.996 ms; incremental P95 is worse. No tail-latency gain is claimed.

Frozen design, failures, paired results and intervals

UR5e final-plan validation

180 constructed roots · 139 unsafe · 41 safe

Six scenario classes compare four policies on the same roots. A separate reviewer uses denser static sampling and 1 ms MuJoCo replay without importing the gate runner.

106 static-prefix reuses and 74 full fallbacks. Incremental validation costs 147.65 ms including the parent, versus 48.74 ms for full validation: no speedup claim. These are fixed-order observations on constructed cases, not a natural-distribution holdout or continuous collision proof.

Methods, denominators and cost accounting

Fresh-state full-suite confirmation: 93/100

10 tasks × states 30–39 · MuJoCo 3.3.7 · execute 10 / predict 50

A 100-rollout confirmation, with source and protocol frozen before execution, succeeds in 93/100 (descriptive 95% Wilson interval 86.25–96.57%), with zero crashes. Seven 280-step failures remain: four in the top-drawer task, one each in tasks 3, 5 and 8. All 11,592 official step outcomes and unchanged environment actions were independently audited. This evaluates the official native Panda checkpoint, separately from the SO100 overlay and Sentinel intervention benefit.

All ten tasks on fresh states 30–39: successes10,10,10,9,6,9,10,10,9,10 and descriptive95% Wilson intervals
The original 58/100 and this 93/100 use different initial states, backend and execution horizon. They are separate configuration results, not a paired improvement estimate. Every failure remains in the fixed 100-cell denominator; the top-drawer task is still 6/10.
Frozen design, failures, paired results and intervals

Original native configuration: 58/100

LIBERO-Spatial · 10 tasks × 10 fixed states

Under MuJoCo 3.8.1 with a 50-action execution horizon, the official SmolVLA LIBERO checkpoint succeeds in 58/100 fixed rollouts (descriptive 95% Wilson interval 48.21–67.20%). The bowl-on-ramekin task fails 10/10; all 42 failures reach the 280-step cap. There are 17,868 unchanged postprocessor-to-env.step actions. This measures the native official checkpoint and logging path, separately from the SO100 overlay and intervention benefit.

Official model native 100-rollout task outcomes and matched 40-rollout feedback interval ablation
Matched feedback ablation, 40 new rollouts: task 5 improves only 0/10 → 1/10; control 8/10 → 10/10. Paired intervals include zero. More frequent inference does not resolve the severe failure.
20 matched fresh-state pairs: task 5 success 0 versus 7 and control 10 versus 10 under complete MuJoCo versions
40 fresh-state rollouts, frozen before execution: task 5 succeeds 0/10 under MuJoCo 3.8.1 versus 7/10 under 3.3.7; control is 10/10 under both. Paired descriptive 95% interval: +40–100 percentage points for task 5. This changes the complete backend version, not a single reset mechanism, and does not establish physical accuracy or overall model success.
Does the simulator version change the task before the first action?

40 no-policy resets isolate MuJoCo 3.8.1 versus 3.3.7 with unchanged task files, assets and paired seeds. Task 5 bowl position differs by 35.327 mm in all 10 pairs; control mean shift is 0.688 mm. This confirms version-sensitive initial conditions, without establishing why the policy failed. The older engine is a benchmark-compatibility comparison, not more accurate physics.

First ordered task0 and task5 reset RGB under the two frozen versions
Frozen design, failures, paired results and intervals

SmolVLA on recorded SO100 actions

Five held-out episodes · 1,926 overlapping windows

A frozen vision-language backbone with a fine-tuned action expert. The pinned base and development-selected overlay use identical sampling noise in one paired held-out evaluation.

Action reconstruction, not physical robot task success. Windows overlap; the independent unit is the episode. Dataset-native physical units are undeclared.

Action MAE / training action standard deviation

63.46% lower normalized MAE

Model card, checkpoint and loading procedure

WorldGuard under distribution shift

600 fresh roots, six scenarios, four sibling plans per root. The trained state and visual ensembles keep their original calibration; difficult cases remain in the comparison.

Unsafe selections / 100 roots · lower is better

Full scenario results and uncertainty

MUJOCO PHYSICS-PROFILE FALLBACK

When every original candidate fails, change the action family.

A separate 5 s MuJoCo profile evaluates 100 new low-friction roots. Its 4.8 s candidate has 0/100 unsafe selections and 100/100 tray endpoint tasks, at three times the 1.6 s duration.

The rule assumes μ_min = 0.015; friction is not sensed. The 3.2 s candidate was also safe in this sample but rejected by the conservative rule. Zero observed failures has a 3.70% Wilson upper bound. This profile is not deployed and does not repair the original neural gate.

Fallback rule and retained results
Measured fixed 1.6 second, long 4.8 second and known-profile fallback comparisons

GPU measurements: RTX 4090 D (24 GB). Every chart uses retained result files, available with source hashes in the evidence download.

Download chart data + source identities

06 / RESEARCH WORKBENCH

Run it locally. Inspect the evidence.

Actual local Sentinel EVC workbench with recorded three-dimensional trajectory and signed evidence

CLI, desktop and SSH entrypoints

The installable CLI and macOS App share a local engine. Run numeric and fixed MuJoCo contact profiles, replay recorded trajectories and export signed evidence bundles.

  • OBJ, STL, MJCF and URDF inspection
  • Strict SSH experiments with source-bound results
  • Read-only robot diagnostics
  • Independent evidence verification

Research prototype. The macOS App is locally built without distribution signing or notarization; imported models do not acquire motion permission. The physical robot port remains read-only.

Installation and supported capabilities

07 / REPRODUCE

From source to a verifiable run.

Local quick start
git clone https://github.com/LancerLSY/sentinel-evc-lab.git
cd sentinel-evc-lab
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[physics]"
python -m sentinel_evc serve

Start from source with Python 3.10+, or use the interactive installer to choose a core or physics profile and build the macOS App. The CLI prints the local workbench address; simulations write recorded outcomes and evidence.

GPU training uses a separate pinned environment. Model overlays require the original frozen SmolVLA base/backbone; they are not standalone policy checkpoints.

Training, calibration and evaluation commands

The historical simulator directory checksum includes download cache. Diagnostic replay source binds official model content per file. Replay reproducibility note

Paper-validation asset index and SHA-256

The release index includes sizes and SHA-256 digests; archives contain per-file identities and licenses. Sentinel code is MIT; MIT does not grant patent rights. See the repository for upstream attribution and the core mechanism's patent notice.

Asset index and checksums

08 / CITATION

Reference the research software.

Use the repository citation for the current software and simulation artifacts. A paper citation will be added alongside its arXiv record.

BibTeX
@misc{sentinel_evc_lab_2026,
  title = {Sentinel EVC Lab},
  year = {2026},
  url = {https://github.com/LancerLSY/sentinel-evc-lab},
  note = {Research software and simulation evidence}
}