Research / Embodied AI safety

SENTINEL-EVC / ROBOTICS

Sentinel-EVC

From the final action to traceable execution.

v0.3 · Public research update, 02 October 2026

Sentinel-EVC places verification at the final action commit point. The new public research adds official SmolVLA closed-loop Panda simulation, version comparisons retaining failures, and support-aware routing with broader physical fallback studies, keeping actions, outcomes, and configurations inspectable.

Original conceptual artwork of a robot, an action trajectory and a validation boundary
Original concept artwork · Recorded experiments and trajectories below

Why must a check keep up with change?

A policy’s action may be retimed, repaired, spliced or transformed between frames. A verdict about the original plan may not cover the motion finally sent out. Sentinel-EVC places the check after transformation and before dispatch, so the decision concerns the version that would execute.

The object of verification is the final action handed to the executor.

Connect the check, the permit and the evidence.

The system asks three questions: does the action pass the current checks, does its permit still match the context, and can an independent reader reconstruct what happened? A changed plan or context returns to rechecking; unsupported candidates remain unknown.

  1. 01

    Check the final version

    Inspect the transformed plan against physical constraints and its bound consequence profile.

  2. 02

    Limit permission in time and use

    Bind a permit to the plan, context and generation, then recheck fresh feedback at submission.

  3. 03

    Retain an inspectable trace

    Distinguish submitted, accepted and observed actions, then replay and verify independently.

Execution chain and recovery rules
Mechanism diagram: upstream action, transform, final checks, expiring one-use permit, single writer, and signed evidence.
Conceptual mechanism diagram; the capability matrix states each profile's actual runtime integration.
  1. 01

    Propose

    A VLA, planner, or recorded policy proposes an action block.

  2. 02

    Transform

    Retiming, repair, and frame conversion produce the plan that would actually execute.

  3. 03

    Recheck

    Physical checks and the bound consequence-prediction profile evaluate the final candidate; unsupported actions remain unknown.

  4. 04

    Permit

    The permit binds plan and context, with a deadline and one use.

  5. 05

    Execute

    The single writer rechecks fresh state and distinguishes submitted, accepted, and observed commands.

  6. 06

    Verify

    Events, results, and asset digests enter signed records for offline replay and verification.

Revocation prevents further old-generation submissions. Recovery confirms cancel and drain, obtains post-revoke feedback, and resumes with the current epoch and explicit approval. Actions already accepted by the controller remain part of the record.

First, let the policy act in its native simulator.

The official LIBERO/Panda policy completes 58/100 in its original configuration. A later fresh-state confirmation uses MuJoCo 3.3.7 and reobserves every ten actions, completing 93/100 across ten tasks while retaining seven timeouts; the top-drawer task remains weakest at 6/10. The runs use different grids/configurations, and both record zero interventions.

93/100

Completed / 100 simulated episodes

Official LIBERO/Panda policy · MuJoCo 3.3.7 · predict 50 / execute 10 · Fresh states 30–39 for each task · Zero observer interventions.

95% Wilson interval: 86.25–96.57%. All seven timeouts remain in the results; the top-drawer task is still 6/10.

Ten tasks, with every denominator retained.
  1. Task 0010/10
  2. Task 0110/10
  3. Task 0210/10
  4. Task 039/10
  5. Task 046/10
  6. Task 059/10
  7. Task 0610/10
  8. Task 0710/10
  9. Task 089/10
  10. Task 0910/10

Official LIBERO/Panda policy · MuJoCo 3.3.7 · predict 50 / execute 10 · Fresh states 30–39 for each task · Zero observer interventions.

A separate grid/backend/horizon from 58/100; zero interventions. It is not a 35-point Sentinel gain, hardware performance, or SO100 task success.

Inspect the original research record
Fresh-state full-suite confirmation
Original Sentinel research figure: full suite confirmation
Original public research record · 02 October 2026

A separate grid/backend/horizon from 58/100; zero interventions. It is not a 35-point Sentinel gain, hardware performance, or SO100 task success.

Complete original-configuration record
Original Sentinel research figure: native policy and horizon
Original public research record · 02 October 2026

All 42 failures are retained; observe-only, distinct from the SO100 overlay, UR5e, and Sentinel intervention efficacy.

Inspect the original research record

Break a failure into testable questions.

Task 5 fails 10/10 originally. More frequent feedback yields 1/10; changing the complete MuJoCo version on fresh paired states yields 0/10 versus 7/10, with the control at 10/10 in both. A no-policy reset audit also finds an initial bowl-pose difference, without isolating one physical cause.

Original Sentinel research figure: backend version success
Complete MuJoCo version comparison · Same official model and predict 50 / execute 10 · Task 5: 0/10 versus 7/10; control task 0: 10/10 in both.
Open the original figure

Success and failure in the original configuration

Successful rollout · Task 0 / state 0 of the original 58/100 configuration · 83 actual actions · 4.15 s; not part of the fresh-state 93/100 cohort.
Failure replay · Task 5 / state 0 of the original 58/100 configuration · 280 actual actions · 14.05 s; a post-hoc exact replay, not an additional experiment.
Horizon, reset and complete-version comparisons

The horizon ablation uses 20 matched states and 40 rollouts; the reset audit compares 40 resets without loading a policy; the complete-version comparison uses a separate 20 matched states and 40 rollouts. They answer different questions and should not be combined into one improvement percentage.

Task 5 still fails 9/10; small-sample intervals include zero. This is not an overall model gain or an asynchronous chunk method.

Inspect the original research record
Original Sentinel research figure: reset version comparison
Original public research record · 02 October 2026

Shows a version difference before policy action; it neither proves reset-pose causation nor ranks physical accuracy.

Inspect the original research record

A complete-version benchmark-compatibility effect, without isolating one contact/reset cause. Task 5 still fails 3/10 under the compatibility version.

Inspect the original research record

Retain task utility alongside unsafe-action rejection.

The new 600-root router completes 300, rejects 300, and observes no unsafe selection, while fixed 4.8 s completes all 600. Across 27 broader stress cells, 4.8 s completes 324/324; 1.6 s completes 184/324 with 78 drops. The floor rule matches fixed slow for every root, at three times the action duration.

Original Sentinel research figure: routing utility
600 fresh simulated cases · Routing completes 300 and rejects 300; fixed 4.8 s completes 600. Rejection counts as incomplete.
Open the original figure
Low-friction replay · First prospectively ordered low-friction case in the new study · Stored simulator states · 10.2 s; the fast action drops, the slow action completes.
27 physical-condition cells and unobserved future friction
Original Sentinel research figure: physical stress
27 stratified condition cells and 324 cases · Fast actions complete 184/324 with 78 drops; fixed slow and the floor rule both complete 324/324, identically on every case.
Open the original figure

The floor rule equals fixed 4.8 s root by root, with no adaptive-duration gain; mu_min=.015 is declared. Finite strata do not qualify a continuous domain or hardware. The 3.2 s analysis is post-hoc.

Inspect the original research record

Across 100 matched cases, observed history, state and action bytes stay identical while future friction changes. Unsafe labels differ in 100/100 pairs for the 1.6 s action and 0/100 for the 4.8 s action.

Shows ambiguity from an unobserved future contact change; it does not prove static friction can never be identified through informative probing.

Inspect the original research record

Declared-profile routing, not automatic pixel OOD or friction estimation. All 300 unsupported mass/displacement rejections count as incomplete; fixed slow has stronger total utility.

Measure rechecking together with its cost.

Full and incremental decisions agree on all 60 fresh obstacle-field roots; parent-only checking falsely allows 20/23 unsafe roots. Fair rotated timing finds a 0.996 ms mean marginal saving with worse P95, supporting no general robot-speed claim.

Original Sentinel research figure: verification cost
60 fresh obstacle-field cases · Three rotated timing repeats with warmup excluded · Full and incremental decisions agree 60/60; mean marginal times are 85.351 ms and 84.356 ms.
Open the original figure

Mean marginal saving: 0.996 ms, paired difference interval −1.728 to −0.291 ms. Total-cost P95 rises from 223.214 ms to 226.218 ms, worse for the incremental path. Only three cases concern collisions; other hazards mainly use designed joint and tracking conditions.

Inspect the original research record
Mechanism film · Rerendered stored trajectories from the earlier 180-case study · 19.95 s. Red rejected proposals are counterfactual and were not dispatched.

Bring the research into an inspectable workbench.

The local workbench brings candidates, run state, stored trajectories and evidence together so a decision’s basis can be inspected. The CLI, macOS app and browser interface share a local engine; the current robot interface provides read-only diagnostics.

Sentinel's local MuJoCo workbench shows saved 3D motion, completed execution, evidence PASS, incomplete scientific acceptance, and slip.
Actual workbench screenshot: completed execution and intact evidence retain the failed scientific gate.
Current modules and walkthrough
Final-action verification

Immutable plan digests, full physical checks, and bounded prefix reuse; mismatched bindings trigger full fallback.

Product core

Consequence prediction

Evaluate four sibling candidates, retain prediction envelopes and model/calibration identities, then select deterministically from the allowed set.

Numeric core; GPU models are separate research

Permits and runtime state

Time-boxed single-use permits, fresh feedback, generation revocation, confirmed cancel, and drain-based recovery.

Product core

Signed evidence and replay

Retain run, model, prediction, and event records; verify independently, filter events, and download ZIP evidence.

Local workbench

3D asset inspection

Import OBJ, STL, MJCF, and URDF; preview parsed geometry and perform bounded MuJoCo compilation and five-step checks.

Inspection and preview; no motion permit

SSH experiment entry

Use an existing OpenSSH host configuration to run source-bound experiments remotely and verify returned evidence and transfer receipts.

Existing alias, key agent, and known_hosts required

Read-only robot diagnostics

Mock diagnostics and Universal Robots model, mode, and safety-status queries; this surface sends no motion commands.

Read-only; physical motion adapter pending

  1. Install the core or physics profile and start the local workbench.
  2. Create or import a scenario and compare geometry and consequence results for four candidates.
  3. Run the experiment and inspect execution completion, evidence integrity, and scientific acceptance separately.
  4. Replay recorded motion and inspect events and the three runtime cursors.
  5. Request stop, drain, and approved recovery through the lifecycle, retaining the actually accepted tail.
  6. Export a signed ZIP and verify it with the independent verifier.

A recorded workbench run shows completed execution, evidence PASS, incomplete scientific acceptance, and slip together. The macOS App is a locally built research shell managed by a Python environment; distribution signing and notarization remain open.

Workbench validation and product records

Earlier studies: mechanisms, models and scenes.

These earlier studies retain their own samples and configurations. Offline SO100 action reconstruction, the 180-case UR5e comparison and WorldGuard scene studies are reported separately from the new native Panda tasks and fresh cost experiment.

Experimental setup, full results and original figure

The UR5e study compares 180 constructed, same-root simulation cases. Independent higher-resolution review labels 139 unsafe and 41 safe; each method faces the same plan changes.

Parent-only checking falsely allows 125 unsafe cases; full final validation and the incremental-with-fallback path allow none in this set. These are results for the constructed simulations, rather than a guarantee across robot settings.

Decisions and the cost of rechecking

False allows / 139 unsafe cases

Parent only
125/139
Full final validation
0/139
Incremental + fallback
0/139

Mean measured total cost

Parent only
101.15 ms
Full final validation
48.74 ms
Incremental + fallback
147.65 ms

The incremental path takes 147.65 ms including the parent check, versus 48.74 ms for full final validation. This result does not establish a speedup.

Experimental setup, full results and original figure

Six scenario classes with 30 roots each produce 180 same-root comparisons. A separately implemented reviewer uses 0.005 rad static sampling and 1 ms MuJoCo steps, labeling 139 roots unsafe and 41 safe.

UR5e same-root comparison
PolicyAllowed / 180False allows / 139 unsafeFalse rejects / 41 safeMean observed wall time
Parent only166/180125/1390/41101.15 ms
Full final validation41/1800/1390/4148.74 ms
Incremental + fallback41/1800/1390/41147.65 ms
Reject transforms30/1800/13911/41No validation call

106 roots reuse the static prefix and 74 fall back fully. The incremental stage averages 46.50 ms plus 101.15 ms for parent validation: 147.65 ms total, slower than full validation's 48.74 ms. Dynamic replay always starts from frame zero.

Seven representative cases, 720p, 33.75 s including titles and final-state holds. The 287 stored joint-state frames run at 1× for 14.35 s. Rendering reruns neither physics nor the gate.

False allows, false rejects, and parent-inclusive validation cost for 180 constructed UR5e roots.
Parent-only falsely allows 125/139; full and incremental both 0/139. Incremental total cost is 147.65 ms versus 48.74 ms full, without a speedup claim.
Original UR5e review results
Data splits, model package and evaluation settings

The SmolVLA study uses official SO100 PickPlace recordings to fine-tune the action expert and related projections, preserving separate training, selection and held-out evaluation splits. The model package, loader and evaluation records are published on Hugging Face.

63.46%

Lower normalized action MAE versus the frozen base model

Five held-out episodes · Recorded-action reconstruction

The metric concerns reconstruction of recorded actions. Episodes are the independent units; overlapping windows are not independent successful tasks.

Data splits, model package and evaluation settings

Fine-tune the action expert and state/action/time projections over a pinned SmolVLA base and frozen SmolVLM2 backbone. Training runs for 5,000 updates, development selects update 3,750, and 99,880,992 parameters are trainable.

Official SO100 PickPlace: 50 episodes, 19,631 frames, 30 Hz, six joint/gripper fields, and top/wrist cameras. The study uses 30 train / 5 dev / 10 cal / 5 test episodes, with normalization fitted on training data only.

One paired held-out evaluation with identical sampling noise lowers normalized MAE by 63.46%. Five held-out episodes produce 1,926 overlapping windows, each predicting a 50×6 action chunk; the independent unit remains the episode.

SmolVLA held-out action reconstruction
PolicyNormalized action MAENative-field action MAE
Pinned base0.67316613.321098
Dev-selected fine-tune0.2459844.47285

The Hub publishes a 21-file overlay, loader, model card, and evaluation record. The pinned original base and backbone are required; one dev-only [1,50,6] example passed a bitwise equality check.

This result measures reconstruction of recorded actions, not robot task success. The data does not declare common physical units; SO100 overlay results cannot be attributed to the separate UR5e simulator study.

RTX 4090 D (24 GB), four threads, batch one: in-memory images through unnormalized actions measure P50 229.26 ms / P95 236.90 ms over 20 synchronized calls. Video reading, transport, and actuation are excluded.

Left: SmolVLA action MAE falls 63.46% over five held-out episodes. Right: matched calibration reduces unsafe selections from 47 to 0 on the same 500-root camera stress set.
Two distinct experiments: recorded-action reconstruction on the left, simulator camera-profile recalibration on the right. The right-side predictions are unchanged and actions shift from 1.2 to 1.6 s.
Six scenarios and their supported scope

WorldGuard tests consequence prediction under distribution shift. Cameras, payload, friction and the supported action range change the conditions for a decision. The study retains failures alongside selection coverage and simpler comparators.

WorldGuard unsafe selections and XY coverage across six fresh scenarios, retaining low-friction and camera failures.
100 fresh roots per scenario. Under camera shift, the image gate selects 7/100 unsafe actions; every original low-friction candidate is unsafe. The 1.35× move is an out-of-contract bare-model diagnostic; the product path rejects all.

A changed camera needs matched calibration.

On the earlier fixed 500-root stress set, shifting cameras with the original calibration produced 47/500 unsafe selections. Matching the dev/cal profile reduced this to 0/500 by selecting the slower 1.6 s action for all roots. Predictions are unchanged, and XY coverage is 94.0%, below the nominal 95%.

Six scenarios and their supported scope

W0/W1/W2 train 36 members across numerical dynamics, contact consequences, visual objects, and real SO100 joint recordings. Stronger simple comparators remain visible: physical identification beats the strict learned numerical model, while adding images increases normalized SO100 joint-prediction error by 10.0%.

Six scenarios, each with 100 fresh roots and four sibling actions per root. Counts below are unsafe selections / 100.
ScenarioFast geometryFixed 1.6 sState WorldGuardImage WorldGuard
Nominal59000
Hidden low friction100100100100
Low mass76000
High mass75000
Camera shift62007
1.35× displacement (bare-model diagnostic)89100

The 1.35× displacement exceeds the action contract. The declared product path returns MODEL_UNKNOWN and rejects all 100 roots; bare-model values describe degradation, not supported-profile authorization.

Camera and preprocessing identities belong with the model, PCA, action family, and calibration. Under hidden low friction, every original candidate is unsafe; prediction or recalibration cannot create a safe action.

WorldGuard scenario report

A safe action must actually exist.

A separate low-friction profile studies 100 fresh simulated cases, extending a 1.6-second action family with slower candidates. The 4.8-second plan completes 100/100 tray tasks with no drops, at three times the action duration.

Outcomes and target-acceleration screening for 1.6, 3.2, and 4.8 s plans in a new low-friction profile.
100 fresh roots, separate 5 s profile, configured μ_min=0.015. The 4.8 s plan completes 100/100 tray tasks with zero drops at 3× duration; 3.2 s is also safe here but the rule rejects it.
Full low-friction action-family results

A separate 5 s physical profile tests 1.6, 3.2, and 4.8 s plans over 100 fresh low-friction roots. The fixed 1.6 s plan has 100/100 unsafe outcomes, 58/100 drops, and 0/100 tray completions; 4.8 s has 0/100 unsafe outcomes, 0/100 drops, and 100/100 tray completions, at three times the action duration.

Screening uses configured μ_min=0.015, without measuring friction. The 3.2 s plan is also safe in this sample but the conservative rule rejects it. Zero/100 still has a 3.70% Wilson upper bound; this is a new, undeployed simulator profile.

Low-friction action-family report

Read, try and reproduce.

Start with the project site, then follow the code, models and experiment records for closer inspection.

Local installation and startup

The interactive installer selects the core or physics profile, creates an isolated environment, and can build the macOS App. Product runtime requires Python 3.10+ and serves only on 127.0.0.1. A PowerShell installation entry exists for Windows; GPU training and loading use a separate pinned environment.

python -m pip install -e ".[physics]"
python -m sentinel_evc serve --data-dir runs/workbench --port 8765
What this release validates

Each study retains commands, source/environment identities, input revisions, model/calibration digests, and per-root or per-episode artifacts. Signatures establish integrity relative to the chosen public key; replay shows stored data without recomputing the experiment.

  • Numeric core, fixed contact, GPU models, UR5e, and LIBERO protocols are separately recorded profiles.
  • Robot entry is read-only; imported-model inspection grants no motion authority.
  • Zero observed failures, prediction coverage, task completion, and physical incidents are distinct metrics.
  • Signatures establish neither sensor honesty, physical execution causation, nor safety certification.
  • ROS 2 motion, physical actuation, mTLS, production PKI, and notarized distribution remain open.

Sources checked: 02 October 2026

GitHub 0087504 · Hugging Face 1fe97cc

Each result retains its own configuration, sample and scope. Code: MIT. Model and data: Apache-2.0. Official UR5e assets retain their original licenses.

Formal research release · 02 October 2026

KEEP IN TOUCH

Continue the conversation about that final step.

Let’s discuss action verification, research replication and future collaboration.

Contact & conversations