LOCAL ONLY ORIGINAL IMMUTABLE ACTION SINK ENFORCED REFERENCE ONLY — NOT DEPLOYMENT-READY NO CLOUD ACTUATION NO TRAINING
Why these boundaries matter

This browser never actuates hardware. It coordinates the part we can help with digitally: contract mismatches, suspicious proposal behavior, documented fixes, and unresolved evidence. The downloaded package can support a separate local test only after an experienced owner deliberately completes the embodiment-specific gates.

STOP TEST

What this helps with

Flag and visualize candidate errors before another costly physical iteration.

Possible controller, normalization, timing, action-dynamics, and missing-evidence problems move into familiar inbox stacks so they can be reviewed one at a time, in safe clusters, with repeatable rules, or by hand. A flag is something to investigate—not proof that it caused instability.

UNCERTAINTY NO EVIDENCE YET Counts stay blank until a local audit or comparison.

Start here · known local fixture

Triage only. No policy code or hardware command runs.
The less-painful loop
  1. 1
    Load evidencepolicy + episode + controller
  2. 2
    Work the inboxkeep uncertainty visible
  3. 3
    Compare same statesdocument → harmonize → try again
What this toy can and cannot tell youOpen the exact boundary
The loop we are trying to make less painful: turn on a robot, learn through a game, train or finetune a policy, find a broken one, use this toy to see what looks wrong, try a documented fix, check whether it still jitters on the same evidence, combine/harmonize the useful episodes, and try again. What this toy can do today: play and record local episodes, make heterogeneous episodes easier to combine, and compare base and candidate policy outputs without forwarding their actions. Training, changing the weights, closed-loop hardware testing, and deciding whether to run the robot still happen outside the toy. What this does: finds digitally observable mismatches and suspicious policy-output behavior that may contribute to instability, and preserves exactly how each item was handled. What it does not do: predict every fall, repair a controller, retrain the weights, reproduce closed-loop contact and balance, or decide that hardware is safe to run.

RESEARCH PREVIEW · EDUCATIONAL DEMO

COLLABORATIVE RESEARCH TOY

Made as a toy on purpose: lower weight, lower ego, and an invitation to inspect, critique, and extend the work together.

Research boundary + technical composition

A local sandbox meant to invite experienced roboticists to play, inspect, critique, and extend the episode-harmonization and pre-deployment stability workflow. It is intentionally presented as an unfinished research toy, not a finished product or authority. Its flags and suggestions are informational hypotheses—not a validated safety device, manufacturer guidance, or permission to deploy. Use it to prepare better questions before testing; independently verify every decision and keep the robot’s normal physical safeguards and on-site supervision.

Technical composition: MuJoCo advances the local task dynamics. Deterministic seeded scene generation adapts ideas from an archived WebGL procedural-universe engine. Live captions combine simulator truth with robust temporal motion evidence. Jiminy follows a GODEL-inspired recent context + grounded task knowledge → response contract, but no GODEL checkpoint, GPT model, hosted LLM, or remote inference service is running.

Embodiments and policy situations that informed this demo

G1 is concrete. Ackermann is external. Everything else stays generic until its exact contract exists.

No tensor-shape shortcuts. Open the scope notes when you need the embodiment and policy-family boundaries.

Open exact embodiment + policy scope

“Broken” here means a candidate that may be unstable, mismatched, or incompletely declared. The browser accepts a local data-only policy package for audit; real inference stays in the owner’s local policy environment through the action-sink bridge.

REFERENCE IMPLEMENTATION

Unitree G1 humanoid

Informed by the reported workflow: a controller first trained with RL, then adapted with offline imitation data and KL regularization; a π0.5/VLA-style checkpoint being integrated with a SONIC/whole-body stack; and demonstrations from a Chinese dataset. The current demo audits that VLA package context. Native multimodal VLA replay is not implemented: images, task text, temporal history, preprocessing, and action chunks do not yet pass through this state-vector/action-vector bridge.

Screen for 29D joint order · .q vs torque/target semantics · normalization · rate/latency · high/low-level command ownership · action drift/jerk · fall-risk proxies
EXTERNAL DEDICATED RULE FAMILY

Ackermann vehicle

Ackermann steering has dedicated controller-semantic checks, but Ackermann is not a registered robot in this LeRobot checkout. It is useful for candidate driving policies only after the exact external vehicle/controller contract is supplied. A matching output width does not establish that steering commands mean the same thing.

Screen for steering angle vs curvature or yaw rate · velocity mode · wheelbase · limits · command frequency · stale observations · stop/watchdog ownership
ADAPTER-REQUIRED FAMILY

Other arms and manipulators

Useful for base-versus-finetuned arm policies when the package declares the embodiment and an experienced owner provides the native controller adapter and safety envelope.

Screen for joint position/velocity/torque vs Cartesian delta · reference frame · joint and gripper channel order · scaling · contact-sensitive discontinuity · controller limits
Policy formats this path is meant to compare: base and finetuned flat-vector checkpoints and RL → offline-imitation adaptations. VLA package/schema audit is useful now; native multimodal VLA replay is not implemented. Execution boundary: static manifests and safe tensor metadata may be uploaded locally; arbitrary policy code is not executed by the web server. The owner runs it locally and sends only proposed actions into the discard-only trace.
LEROBOT-SHAPED CORE The audit is not tied to 29 G1 joints.

The local LeRobot checkout defines a common robot boundary around observation_features, action_features, get_observation, and send_action. Shot‑5.1.1 audits and records that boundary; it does not assume that two equally sized vectors share physical meaning.

ROBOTS ALREADY REPRESENTED LOCALLY One dedicated reference; the remaining registered embodiments are generic-only

G1 is the concrete LeRobot reference. The other registered implementations—SO followers and bi‑SO, Koch, OMX, OpenArm and bi‑OpenArm, ReBot B601 and bi‑ReBot, Reachy2, HopeJr, LeKiwi, and EarthRover Mini Plus—receive generic package/schema and flat-vector shadow analysis only. Each needs its own Shot‑5.1.1 manifest, semantic rule set, adapter, and safety envelope before meaningful native replay or actuation.

POLICY FAMILIES ALREADY REPRESENTED LOCALLY Behavior cloning, RL, diffusion, and VLA candidates

The policy factory includes ACT, Diffusion, VQ‑BeT, TD‑MPC, Gaussian Actor, π0, π0.5, SmolVLA, GR00T, X‑VLA, and related families. The present bridge compares finite flat action vectors produced from state-vector observations. Native multimodal VLA replay is absent; π0.5, SmolVLA, and similar packages remain package/schema-audit only until a dedicated adapter preserves images, task text, preprocessing, temporal history, and action chunks.

ARCHIVED RESEARCH CONTRIBUTION Seeded task coverage + robust temporal evidence

The procedural-world code contributes deterministic seed streams for reproducible scenario variation. The temporal-evidence code contributes robust median/MAD motion scoring, provisional events, confirmation/retraction, and source attribution. Those methods transfer across embodiments; task geometry, simulator truth, goal predicates, and safety limits do not.

Implemented interpretation today: Dedicated audit rules exist for exactly two families: Unitree G1 and Ackermann steering. G1 is the concrete LeRobot reference; Ackermann is an external controller family, not a robot registered in this checkout. No other embodiment-aware rule set is implemented today. All other registered LeRobot embodiments receive generic package/schema and flat-vector shadow analysis only until their own semantic rules, manifest, and adapter are added. Native multimodal VLA replay is not implemented.

POLICYWeights + inference behavior
POLICY CONTRACTInputs, outputs, normalization + timing
CONTROLLER MANIFESTMapping, ownership, limits + safe stop
EPISODE / SHADOW TRACERecorded evidence of what happened or was proposed

One traceable path

Use recorded evidence to find what could destabilize this finetuned policy before the robot has to discover it physically.

Policy + dataset + controller contract → grouped flag inbox → matched base/candidate shadow traces → action-risk comparison → documented local handoff. The original policies remain unchanged; uncertainty stays visible, and unresolved physical meaning is handed to the domain expert instead of being silently guessed.

First local trial · no arbitrary policy code in the web process

See the G1 failure-screening workflow, then replace the fixture with your friend’s policy

The first example shows how a Chinese/offline dataset, a finetuned policy, and the real G1 controller can disagree in ways that are tedious to catch by hand. The paired shadow demo then exposes suspicious changes on matched observations without moving a robot. Replace those fixtures with the real package, recorded states, and controller manifest to make the inbox specific to the next planned test.

  1. 1Policy ZIPreplace fixture weights + policy contract
  2. 2Recorded episodereplace the registered target collection
  3. 3Exact controller manifestreplace every unknown physical declaration
  4. 4Optional shadow traceadd discarded proposed-action evidence
Output Grouped risk inbox + paired trace comparison + downloadable handoff receipt The paired fixture demonstrates comparison only. It is not a policy to deploy.

Step 01 · policy intake and contract audit

Fill the Policy · Target · Controller Satchel, then triage every mismatch

STOP TEST · AUDIT REQUIRED

Policy · target · controller Satchel

Select the combinations you want to compare

Each checked policy × target pair gets its own receipt. The browser runs bounded local audits one at a time so a large selection cannot overwhelm the service.

0 comparisons
1 Policy Satchel
0 packages
Policies will appear here.
2 Target Satchel
0 episode targets
Episode collections will appear here.
3 Controller manifest

A manifest gives dimensions physical meaning: embodiment, channel order, mode, units, frame, timing, latency, limits, and stale-command behavior.

Manifest required for meaningful controller replay Generic package/schema audit can still run without one.
Inspect parsed manifest
No manifest loaded.
4 Optional shadow evidence

A shadow trace records observations and proposed commands while actions remain discarded. It is evidence—not a policy or controller contract.

No extra trace attached Selected episode targets still provide recorded evidence.

Select at least one policy and one target. The current cap is 12 local comparisons per run.

Add one policy package to the Satchel

ZIP only. Static files are inventoried and hashed; package code is never executed.

loopback upload
No file leaves this computer.

Focused comparison and replay pair

The dropdowns choose one pair for detailed replay. The checkbox Satchel above audits several combinations without mixing their receipts.

Recorded dataset view 100D raw observation vector
Policy input contract 32 named policy inputs
Policy output contract 29D G1 joint-position targets
29D does not mean the same controller.

A dimension match may be green while action mode remains blocked: the simulation declares 29 joint torques, direct G1 hardware exposes 29 .q positions, and a SONIC whole-body controller requires its own declared interface.

MODE BLOCKER
REAL POLICY WORKFLOW

Import your friend’s actual package and review its actual evidence. The bundled incomplete package is only a diagnostic crash-test sample for demonstrating the inbox—not a policy to hand to a robot owner.

An audit may propose an adapter receipt. It cannot make an undeclared unit, frame, timing, or normalization assumption safe.

CONTRACT FLAG INBOX Unresolved semantics wait here
Offline blockers
Warnings
Checks passed
Hardware blockers
Resolution plans0
EXACT RULE TRACE STATISTIC RECORDED HUMAN Every finding states what kind of evidence produced it.
0 selected
Run a contract audit Timing, units, frame, normalizer, input, and action mismatches will appear as familiar inbox threads.

Step 02 · local policy shadow bridge

Help your friend test the broken policy before the robot has to.

ZERO-FORWARD GUARANTEE

You work the toy; your friend keeps the real policy where it already runs.

You can be the remote collaborator working this toy and the shared flag inbox through the human collaboration channel you already use; the toy itself stays loopback-only. Your friend is onsite, currently finetuning manually, and runs base and candidate inside the existing local policy environment. Recorded observations go in; proposed actions come back only as evidence and are discarded. The robot stays still for this round.

Mission: Use every digitally available clue before another physical run. The toy may reduce avoidable stability-testing iterations by surfacing clues related to jitter, falls, or part breakage; it cannot guarantee that the policy is stable or that hardware will not be damaged.

YOUChoose recorded observations
FRIEND’S MACHINERun base + candidate locally
ROBOT STAYS STILLDiscard every proposal
SHARED INBOXReview flags + trace evidence
Forwarded
0
Hardware commands
0
Simulator commands
0
Network egress
0
Policy code in web server
0

Choose the same observation trace for both policies

A fair comparison holds observations, ordering, timestamps, adapter declarations, and frame limits constant.

read-only source
or import a local trace

Nothing leaves this computer. Imported files are treated as data, never code.

Local policy identities · no weights uploaded

Connect the friend’s own local policy environment

The bridge kit speaks a small localhost JSONL protocol. It does not install the policy, resolve its dependencies, import actuator SDKs, or call a hosted model.

create sessions first
  1. 1
    Download the bridge kitOne small client, protocol schema, README, and session receipt template.
  2. 2
    Start it inside the policy environmentGive it a local inference callable; keep controller and actuator processes stopped.
  3. 3
    Run base, then candidateBoth receive the same recorded observations and may only return proposals.
LOCAL COMMAND TEMPLATE
Create a session pair to receive a loopback command.
Download local bridge kit Inspect the kit before running it. The session secret authorizes trace submission only—never actuation.
A
BASE POLICY

Waiting for a session

NOT CREATED
0 / 0 observationsNo local runner connected
Received
0
Discarded
0
Forwarded
0
Hardware
0
Simulator
0
Network
0
Trace receipt
B
CANDIDATE / FINETUNED POLICY

Waiting for a session

NOT CREATED
0 / 0 observationsNo local runner connected
Received
0
Discarded
0
Forwarded
0
Hardware
0
Simulator
0
Network
0
Trace receipt

Choose one observation source, then create two discard-only sessions.

Step 03 · matched-observation comparison

Ask what changed between base and finetuned—without calling it stable

WAITING FOR TWO TRACES

Base ↔ candidate evidence

These flags identify mismatches or suspicious policy behavior that may affect stability. Review and resolve what you can before another physical run; the on-site tester makes the final call about whether and how to proceed.

Matched observationswaiting for traces
Mean output driftnot measured
Candidate jerknot measured
Candidate saturationnot measured
Latency deltanot measured
Policy KL divergencenot_computablerequires comparable action distributions—not point proposals
No comparison yet

Finalize both traces. “Not computable” will remain visible wherever the submitted evidence cannot support a metric.

No hardware command is needed—or reachable—to make this comparison.

Supporting check · bundled offline fixture replay

Keep the frozen Shot‑5 replay path for reproducible contract tests

ACTION SINK ONLY

Offline replay

Read observations, inspect policy outputs, and measure contract-level risk without actuating hardware.

Policy
not selected
Collection
shot1-g1
Resolution plans
0
Execution mode
action_sink_only

Audit a package before replay.

Possible stability risks to review before testing

These flags identify mismatches or suspicious policy behavior that may affect stability. Review and resolve what you can before another physical run; the on-site tester makes the final call about whether and how to proceed.

awaiting replay
Frames inspectedno run
Saturationno run
Command jerkno run
Contract violationsno run
Recorded episode language + replay evidence Parquet language streams · recorded with the episode · not generated by replay
0 frames
Run a replay to inspect event-aligned output behavior.

Step 04 · auditable export and physical return loop

Download evidence and adapters; keep deployment decisions local

NOT HARDWARE APPROVED
Not approved does not erase the domain expert’s autonomy.

This tool does not certify safety or authorize a physical run. After reviewing the unresolved risks and scientific tradeoffs, an experienced robot owner may still independently choose a carefully staged, on-site supervised test at their own risk. That decision remains visibly recorded as the tester’s decision—not converted into a green “safe” state.

  1. TOY VERDICTNOT HARDWARE APPROVED
  2. OWNER CHOICEExperienced onsite tester may independently stage a test at own risk
  3. HANDOFF TRAVELS WITHUnresolved flags · abstentions · changes · rollback · safeguards
  4. WARNING PERSISTSContinuing never turns this green or certified
The browser and server cannot actuate. A downloaded package can reach actuators only after explicit local deployment, an embodiment-specific hardware adapter, deliberate local enablement, verified limits and watchdogs, a physical emergency stop, and an experienced on-site supervisor.

Two downloads with different meanings

The unchanged original and the generated evidence bundle are kept separate by design.

REFERENCE DEFAULT · LOCAL DEPLOYMENT REQUIRES EXPERT ENABLEMENT

Reference handoff package for the friend

A versioned, auditable package for continuing the stability-testing loop locally. It never presents rewritten weights or automatic permission to deploy; dry-run is the default.

  • Unchanged policy copies by hashonly when a Policy Satchel package was explicitly linked
  • Documented overlaysadapter, normalizer, joint order, timing, and controller declarations
  • Review historyall flags, individual or bulk decisions, rules, uncertainty, and abstentions
  • Before / after evidencebase and candidate discard-only traces plus matched-trajectory comparison
  • Rollback manifestsource hashes, versioned artifacts, change log, and parent lineage
  • START HERE overviewpurpose, results, unresolved risks, local gates, supervised test checklist, and return-loop instructions
NOT HARDWARE APPROVED · the domain expert makes and records the final scientific tradeoff.
ORIGINAL PACKAGE Byte-for-byte unchanged

The exact imported ZIP, identified by its SHA‑256 hash. No adapter, declaration, or metric is written into it.

Download original
DOCUMENTED LOCAL HANDOFF START HERE + receipt + evidence + gates + rollback

A companion ZIP containing an overview, hashes, contract diff, adapter and controller overlays, human decisions, explicit abstentions, comparison evidence, local deployment gate, change log, and rollback manifest.

Prepare supervised test-at-your-own-risk handoff
Documented handoff contents START_HERE.md manifest.json comparison.json documented_changes.json flags_and_abstentions.json local_deployment_gate.json

No new checkpoint is produced because Shot‑5.1.1 does not train. The browser bridge records proposals but never forwards them. After download, the experienced owner may install an embodiment-specific local adapter and explicitly enable a guarded actuator path on their own machine. A generated mapping remains an interface transformation—not evidence that the weights are stable or that a run is approved.

BETA · GUARDED HARDWARE RETURN LOOP

Semi-automates the tedious stability-testing cycle around policy finetuning: schema harmonization, contract validation, matched-observation shadow inference, action-drift and anomaly analysis, failure clustering, human review, versioned handoff, and rollback. After a supervised physical run, the robot owner returns the recorded episode and safety trace so measured outcomes can be compared with pre-run predictions. The tool never promotes a checkpoint automatically; the on-site tester reviews the evidence and makes every physical deployment decision.

  1. 1
    Dry-run locallyverify hashes, dimensions, normalizers, timing, and command bounds with hardware disconnected
    first gate
  2. 2
    Shadow commandsread live observations and log proposed actions without passing them to actuators
    local only
  3. 3
    Expert-decided guarded hardware testthe domain expert weighs unresolved risk against scientific value; independent safety controller, physical e-stop, conservative limits, and trained operator remain mandatory
    at own risk
  4. 4
    return the evidenceupload the resulting episode and controller logs to the episode harmonizer for comparison
    close loop
NOT HARDWARE APPROVED—even when the owner chooses to continue.

The downloaded handoff supports a separate, explicit local decision; it does not turn that decision into certification. Offline replay covers recorded states only, and a policy can still diverge under contact, latency, sensor faults, or out-of-distribution states.

open the shot‑5.1.1 episode harmonizer for returned runs →