01 / THE FINAL HANDOFF
Why must a check keep up with change?
A policy’s action may be retimed, repaired, spliced or transformed between frames. A verdict about the original plan may not cover the motion finally sent out. Sentinel-EVC places the check after transformation and before dispatch, so the decision concerns the version that would execute.
The object of verification is the final action handed to the executor.
02 / CHECK · PERMIT · TRACE
Connect the check, the permit and the evidence.
The system asks three questions: does the action pass the current checks, does its permit still match the context, and can an independent reader reconstruct what happened? A changed plan or context returns to rechecking; unsupported candidates remain unknown.
- 01
Check the final version
Inspect the transformed plan against physical constraints and its bound consequence profile.
- 02
Limit permission in time and use
Bind a permit to the plan, context and generation, then recheck fresh feedback at submission.
- 03
Retain an inspectable trace
Distinguish submitted, accepted and observed actions, then replay and verify independently.
Execution chain and recovery rules
- 01
Propose
A VLA, planner, or recorded policy proposes an action block.
- 02
Transform
Retiming, repair, and frame conversion produce the plan that would actually execute.
- 03
Recheck
Physical checks and the bound consequence-prediction profile evaluate the final candidate; unsupported actions remain unknown.
- 04
Permit
The permit binds plan and context, with a deadline and one use.
- 05
Execute
The single writer rechecks fresh state and distinguishes submitted, accepted, and observed commands.
- 06
Verify
Events, results, and asset digests enter signed records for offline replay and verification.
Revocation prevents further old-generation submissions. Recovery confirms cancel and drain, obtains post-revoke feedback, and resumes with the current epoch and explicit approval. Actions already accepted by the controller remain part of the record.
03 / NATIVE POLICY · NEW STATES
First, let the policy act in its native simulator.
The official LIBERO/Panda policy completes 58/100 in its original configuration. A later fresh-state confirmation uses MuJoCo 3.3.7 and reobserves every ten actions, completing 93/100 across ten tasks while retaining seven timeouts; the top-drawer task remains weakest at 6/10. The runs use different grids/configurations, and both record zero interventions.
Completed / 100 simulated episodes
Official LIBERO/Panda policy · MuJoCo 3.3.7 · predict 50 / execute 10 · Fresh states 30–39 for each task · Zero observer interventions.
95% Wilson interval: 86.25–96.57%. All seven timeouts remain in the results; the top-drawer task is still 6/10.
- Task 0010/10
- Task 0110/10
- Task 0210/10
- Task 039/10
- Task 046/10
- Task 059/10
- Task 0610/10
- Task 0710/10
- Task 089/10
- Task 0910/10
Official LIBERO/Panda policy · MuJoCo 3.3.7 · predict 50 / execute 10 · Fresh states 30–39 for each task · Zero observer interventions.
A separate grid/backend/horizon from 58/100; zero interventions. It is not a 35-point Sentinel gain, hardware performance, or SO100 task success.
Inspect the original research recordFresh-state full-suite confirmation
A separate grid/backend/horizon from 58/100; zero interventions. It is not a 35-point Sentinel gain, hardware performance, or SO100 task success.
Complete original-configuration record
All 42 failures are retained; observe-only, distinct from the SO100 overlay, UR5e, and Sentinel intervention efficacy.
Inspect the original research record04 / FAILURE · COMPARISON
Break a failure into testable questions.
Task 5 fails 10/10 originally. More frequent feedback yields 1/10; changing the complete MuJoCo version on fresh paired states yields 0/10 versus 7/10, with the control at 10/10 in both. A no-policy reset audit also finds an initial bowl-pose difference, without isolating one physical cause.
Success and failure in the original configuration
Horizon, reset and complete-version comparisons
The horizon ablation uses 20 matched states and 40 rollouts; the reset audit compares 40 resets without loading a policy; the complete-version comparison uses a separate 20 matched states and 40 rollouts. They answer different questions and should not be combined into one improvement percentage.
Task 5 still fails 9/10; small-sample intervals include zero. This is not an overall model gain or an asynchronous chunk method.
Inspect the original research record
Shows a version difference before policy action; it neither proves reset-pose causation nor ranks physical accuracy.
Inspect the original research recordA complete-version benchmark-compatibility effect, without isolating one contact/reset cause. Task 5 still fails 3/10 under the compatibility version.
Inspect the original research record05 / UTILITY · PHYSICAL FALLBACK
Retain task utility alongside unsafe-action rejection.
The new 600-root router completes 300, rejects 300, and observes no unsafe selection, while fixed 4.8 s completes all 600. Across 27 broader stress cells, 4.8 s completes 324/324; 1.6 s completes 184/324 with 78 drops. The floor rule matches fixed slow for every root, at three times the action duration.
27 physical-condition cells and unobserved future friction
The floor rule equals fixed 4.8 s root by root, with no adaptive-duration gain; mu_min=.015 is declared. Finite strata do not qualify a continuous domain or hardware. The 3.2 s analysis is post-hoc.
Inspect the original research recordAcross 100 matched cases, observed history, state and action bytes stay identical while future friction changes. Unsafe labels differ in 100/100 pairs for the 1.6 s action and 0/100 for the 4.8 s action.
Shows ambiguity from an unobserved future contact change; it does not prove static friction can never be identified through informative probing.
Inspect the original research recordDeclared-profile routing, not automatic pixel OOD or friction estimation. All 300 unsupported mass/displacement rejections count as incomplete; fixed slow has stronger total utility.
06 / FINAL CHECK · MEASURED COST
Measure rechecking together with its cost.
Full and incremental decisions agree on all 60 fresh obstacle-field roots; parent-only checking falsely allows 20/23 unsafe roots. Fair rotated timing finds a 0.996 ms mean marginal saving with worse P95, supporting no general robot-speed claim.
Mean marginal saving: 0.996 ms, paired difference interval −1.728 to −0.291 ms. Total-cost P95 rises from 223.214 ms to 226.218 ms, worse for the incremental path. Only three cases concern collisions; other hazards mainly use designed joint and tracking conditions.
Inspect the original research record07 / RESEARCH INTO PRACTICE
Bring the research into an inspectable workbench.
The local workbench brings candidates, run state, stored trajectories and evidence together so a decision’s basis can be inspected. The CLI, macOS app and browser interface share a local engine; the current robot interface provides read-only diagnostics.

Current modules and walkthrough
- Final-action verification
Immutable plan digests, full physical checks, and bounded prefix reuse; mismatched bindings trigger full fallback.
Product core
- Consequence prediction
Evaluate four sibling candidates, retain prediction envelopes and model/calibration identities, then select deterministically from the allowed set.
Numeric core; GPU models are separate research
- Permits and runtime state
Time-boxed single-use permits, fresh feedback, generation revocation, confirmed cancel, and drain-based recovery.
Product core
- Signed evidence and replay
Retain run, model, prediction, and event records; verify independently, filter events, and download ZIP evidence.
Local workbench
- 3D asset inspection
Import OBJ, STL, MJCF, and URDF; preview parsed geometry and perform bounded MuJoCo compilation and five-step checks.
Inspection and preview; no motion permit
- SSH experiment entry
Use an existing OpenSSH host configuration to run source-bound experiments remotely and verify returned evidence and transfer receipts.
Existing alias, key agent, and known_hosts required
- Read-only robot diagnostics
Mock diagnostics and Universal Robots model, mode, and safety-status queries; this surface sends no motion commands.
Read-only; physical motion adapter pending
- Install the core or physics profile and start the local workbench.
- Create or import a scenario and compare geometry and consequence results for four candidates.
- Run the experiment and inspect execution completion, evidence integrity, and scientific acceptance separately.
- Replay recorded motion and inspect events and the three runtime cursors.
- Request stop, drain, and approved recovery through the lifecycle, retaining the actually accepted tail.
- Export a signed ZIP and verify it with the independent verifier.
A recorded workbench run shows completed execution, evidence PASS, incomplete scientific acceptance, and slip together. The macOS App is a locally built research shell managed by a Python environment; distribution signing and notarization remain open.
Workbench validation and product records08 / HISTORICAL RECORDS
Earlier studies: mechanisms, models and scenes.
These earlier studies retain their own samples and configurations. Offline SO100 action reconstruction, the 180-case UR5e comparison and WorldGuard scene studies are reported separately from the new native Panda tasks and fresh cost experiment.
Experimental setup, full results and original figure
The UR5e study compares 180 constructed, same-root simulation cases. Independent higher-resolution review labels 139 unsafe and 41 safe; each method faces the same plan changes.
Parent-only checking falsely allows 125 unsafe cases; full final validation and the incremental-with-fallback path allow none in this set. These are results for the constructed simulations, rather than a guarantee across robot settings.
False allows / 139 unsafe cases
- Parent only
- 125/139
- Full final validation
- 0/139
- Incremental + fallback
- 0/139
Mean measured total cost
- Parent only
- 101.15 ms
- Full final validation
- 48.74 ms
- Incremental + fallback
- 147.65 ms
The incremental path takes 147.65 ms including the parent check, versus 48.74 ms for full final validation. This result does not establish a speedup.
Experimental setup, full results and original figure
Six scenario classes with 30 roots each produce 180 same-root comparisons. A separately implemented reviewer uses 0.005 rad static sampling and 1 ms MuJoCo steps, labeling 139 roots unsafe and 41 safe.
| Policy | Allowed / 180 | False allows / 139 unsafe | False rejects / 41 safe | Mean observed wall time |
|---|---|---|---|---|
| Parent only | 166/180 | 125/139 | 0/41 | 101.15 ms |
| Full final validation | 41/180 | 0/139 | 0/41 | 48.74 ms |
| Incremental + fallback | 41/180 | 0/139 | 0/41 | 147.65 ms |
| Reject transforms | 30/180 | 0/139 | 11/41 | No validation call |
106 roots reuse the static prefix and 74 fall back fully. The incremental stage averages 46.50 ms plus 101.15 ms for parent validation: 147.65 ms total, slower than full validation's 48.74 ms. Dynamic replay always starts from frame zero.
Seven representative cases, 720p, 33.75 s including titles and final-state holds. The 287 stored joint-state frames run at 1× for 14.35 s. Rendering reruns neither physics nor the gate.
Data splits, model package and evaluation settings
The SmolVLA study uses official SO100 PickPlace recordings to fine-tune the action expert and related projections, preserving separate training, selection and held-out evaluation splits. The model package, loader and evaluation records are published on Hugging Face.
Lower normalized action MAE versus the frozen base model
Five held-out episodes · Recorded-action reconstruction
The metric concerns reconstruction of recorded actions. Episodes are the independent units; overlapping windows are not independent successful tasks.
Data splits, model package and evaluation settings
Fine-tune the action expert and state/action/time projections over a pinned SmolVLA base and frozen SmolVLM2 backbone. Training runs for 5,000 updates, development selects update 3,750, and 99,880,992 parameters are trainable.
Official SO100 PickPlace: 50 episodes, 19,631 frames, 30 Hz, six joint/gripper fields, and top/wrist cameras. The study uses 30 train / 5 dev / 10 cal / 5 test episodes, with normalization fitted on training data only.
One paired held-out evaluation with identical sampling noise lowers normalized MAE by 63.46%. Five held-out episodes produce 1,926 overlapping windows, each predicting a 50×6 action chunk; the independent unit remains the episode.
| Policy | Normalized action MAE | Native-field action MAE |
|---|---|---|
| Pinned base | 0.673166 | 13.321098 |
| Dev-selected fine-tune | 0.245984 | 4.47285 |
The Hub publishes a 21-file overlay, loader, model card, and evaluation record. The pinned original base and backbone are required; one dev-only [1,50,6] example passed a bitwise equality check.
This result measures reconstruction of recorded actions, not robot task success. The data does not declare common physical units; SO100 overlay results cannot be attributed to the separate UR5e simulator study.
RTX 4090 D (24 GB), four threads, batch one: in-memory images through unnormalized actions measure P50 229.26 ms / P95 236.90 ms over 20 synchronized calls. Video reading, transport, and actuation are excluded.
Six scenarios and their supported scope
WorldGuard tests consequence prediction under distribution shift. Cameras, payload, friction and the supported action range change the conditions for a decision. The study retains failures alongside selection coverage and simpler comparators.
A changed camera needs matched calibration.
On the earlier fixed 500-root stress set, shifting cameras with the original calibration produced 47/500 unsafe selections. Matching the dev/cal profile reduced this to 0/500 by selecting the slower 1.6 s action for all roots. Predictions are unchanged, and XY coverage is 94.0%, below the nominal 95%.
Six scenarios and their supported scope
W0/W1/W2 train 36 members across numerical dynamics, contact consequences, visual objects, and real SO100 joint recordings. Stronger simple comparators remain visible: physical identification beats the strict learned numerical model, while adding images increases normalized SO100 joint-prediction error by 10.0%.
| Scenario | Fast geometry | Fixed 1.6 s | State WorldGuard | Image WorldGuard |
|---|---|---|---|---|
| Nominal | 59 | 0 | 0 | 0 |
| Hidden low friction | 100 | 100 | 100 | 100 |
| Low mass | 76 | 0 | 0 | 0 |
| High mass | 75 | 0 | 0 | 0 |
| Camera shift | 62 | 0 | 0 | 7 |
| 1.35× displacement (bare-model diagnostic) | 89 | 1 | 0 | 0 |
The 1.35× displacement exceeds the action contract. The declared product path returns MODEL_UNKNOWN and rejects all 100 roots; bare-model values describe degradation, not supported-profile authorization.
Camera and preprocessing identities belong with the model, PCA, action family, and calibration. Under hidden low friction, every original candidate is unsafe; prediction or recalibration cannot create a safe action.
WorldGuard scenario reportA safe action must actually exist.
A separate low-friction profile studies 100 fresh simulated cases, extending a 1.6-second action family with slower candidates. The 4.8-second plan completes 100/100 tray tasks with no drops, at three times the action duration.
Full low-friction action-family results
A separate 5 s physical profile tests 1.6, 3.2, and 4.8 s plans over 100 fresh low-friction roots. The fixed 1.6 s plan has 100/100 unsafe outcomes, 58/100 drops, and 0/100 tray completions; 4.8 s has 0/100 unsafe outcomes, 0/100 drops, and 100/100 tray completions, at three times the action duration.
Screening uses configured μ_min=0.015, without measuring friction. The 3.2 s plan is also safe in this sample but the conservative rule rejects it. Zero/100 still has a 3.70% Wilson upper bound; this is a new, undeployed simulator profile.
Low-friction action-family report09 / EXPLORE THE PROJECT
Read, try and reproduce.
Start with the project site, then follow the code, models and experiment records for closer inspection.
Experiments & records
Simulation research releaseFile sizes and SHA-256 indexOriginal GPU research releaseLocal installation and startup
The interactive installer selects the core or physics profile, creates an isolated environment, and can build the macOS App. Product runtime requires Python 3.10+ and serves only on 127.0.0.1. A PowerShell installation entry exists for Windows; GPU training and loading use a separate pinned environment.
python -m pip install -e ".[physics]"
python -m sentinel_evc serve --data-dir runs/workbench --port 8765What this release validates
Each study retains commands, source/environment identities, input revisions, model/calibration digests, and per-root or per-episode artifacts. Signatures establish integrity relative to the chosen public key; replay shows stored data without recomputing the experiment.
- Numeric core, fixed contact, GPU models, UR5e, and LIBERO protocols are separately recorded profiles.
- Robot entry is read-only; imported-model inspection grants no motion authority.
- Zero observed failures, prediction coverage, task completion, and physical incidents are distinct metrics.
- Signatures establish neither sensor honesty, physical execution causation, nor safety certification.
- ROS 2 motion, physical actuation, mTLS, production PKI, and notarized distribution remain open.
Sources checked: 02 October 2026
GitHub 0087504 · Hugging Face 1fe97cc
Each result retains its own configuration, sample and scope. Code: MIT. Model and data: Apache-2.0. Official UR5e assets retain their original licenses.
