Haga

Data room

External investor sharing

Evidence

Evidence snapshots

Ready

Curated numeric snapshot of current public metrics — narrative companion to live charts.

Living narrative of what the public metrics API currently shows. Live charts stay on the Evidence index (same publish channel as the public Lab / landing). Raw episode dumps never appear in this room.

Figures below mirror fixtures / published artifacts as of 2026-07-17. Prefer the charts if schema dates diverge.


Claim boundaries (read first)

Bound What we claim What we do not claim
Scope Sim-first MuJoCo / Robosuite + position-only / tracked-video physics checks Real-hardware validation
Policies Scripted OSC_POSE baselines on Lift, Stack, PickPlaceCan, Door Exhaustive multi-policy leaderboard
Gate vs stress Mild is the gate; degradation under moderate/severe is the story That mild PASS alone proves robustness
Door gate Success-primary (grasp N/A for rotating handle) Same grasp-fail semantics as Lift/Stack
Generative CogVideoX-5b-I2V cohort, n=6, seeds 0–1, static_hover Cosmos / Genie / NIM scores or broad generative quality
Ecology Pre-deployment stress check Substitute for Robocurve / RoboArena

Progress digest

Field Value
Phase 3 — Public methodology
Status Current (article + CTA live)
Next milestone 2–3 live design-partner evaluation reports
Last evidence 2026-07-17

Day-to-day GitHub Issues are not mirrored here.


Pillar 1 — dual-task degradation summary

Task Nominal → severe Mild gate Severe failure mode Artifact
Lift 1.00 → 0.26 PASS Grasp slip 0.74 pillar1-lift.json
Stack 0.96 → 0.20 PASS Grasp fail 0.80 pillar1-stack.json
PickPlaceCan 1.00 → 0.24 PASS Grasp slip 0.76 pillar1-pickplacecan.json
Door 0.60 → 0.52 PASS (success-primary) Door did not open 0.48 pillar1-door.json

Setup (all four): 50 episodes per condition · paired seeds 0..49 · Wilson 95% CIs · Panda · OSC_POSE. Headline rates from published metrics JSON (Lab / Evidence charts).

Pillar 1 — Lift policy stress (Panda)

Setup: 50 episodes per condition · paired seeds · Wilson 95% CIs · gate tier = mild.

Condition Success Success CI Grasp-fail Peak EE force (mean) Verdict
Nominal 1.00 [0.93, 1.00] 0.00 7.30
Mild 1.00 [0.93, 1.00] 0.00 7.82 PASS
Moderate 0.90 [0.79, 0.96] 0.10 12.77 PASS
Severe 0.26 [0.16, 0.40] 0.74 14.15 FAIL

Headline: success degrades 1.00 → 0.26 from nominal/mild to severe; under severe stress the gripper acquires the cube then loses grasp (slip) in 74% of episodes.


Pillar 1 — Stack policy stress (Panda)

Setup: same tier model · 50 episodes per condition · scripted OSC stack baseline · gate = mild.

Condition Success Success CI Grasp-fail Peak EE force (mean) Verdict
Nominal 0.96 [0.87, 0.99] 0.04 7.44
Mild 0.82 [0.69, 0.90] 0.18 8.35 PASS
Moderate 0.72 [0.58, 0.83] 0.28 12.98 FAIL
Severe 0.20 [0.11, 0.33] 0.80 13.58 FAIL

Headline: success degrades 0.96 → 0.20; severe grasp-fail 0.80.


Pillar 1 — PickPlaceCan policy stress (Panda)

Setup: same tier model · 50 episodes per condition · gate = mild.

Condition Success Success CI Grasp-fail Peak EE force (mean) Verdict
Nominal 1.00 [0.93, 1.00] 0.00 165.86
Mild 1.00 [0.93, 1.00] 0.00 214.54 PASS
Moderate 0.88 [0.76, 0.94] 0.12 160.43 PASS
Severe 0.24 [0.14, 0.37] 0.76 46.87 FAIL

Headline: success degrades 1.00 → 0.24; severe grasp-slip 0.76.


Pillar 1 — Door policy stress (Panda)

Setup: panel-scaled mass/friction tiers · success-primary gate (grasp N/A for rotating handle) · 50 episodes per condition.

Condition Success Success CI Grasp-fail Peak EE force (mean) Verdict
Nominal 0.60 [0.46, 0.72] 0.00 (N/A) 209.02
Mild 0.64 [0.50, 0.76] 0.00 (N/A) 185.57 PASS
Moderate 0.58 [0.44, 0.71] 0.00 (N/A) 199.09 PASS
Severe 0.52 [0.39, 0.65] 0.00 (N/A) 191.91 PASS

Headline: success 0.60 → 0.52 — shallower curve than grasp tasks; documented failure mode is door did not open. Gate PASS does not erase the ~48% severe fail rate.

How to read Pillar 1: mild-gate PASS is expected — the diligence story is the degradation curve and documented failure modes across tasks.


Pillar 2 — Physics-consistency checker calibration

Setup: MuJoCo ground-truth trajectories with labeled violations · position-only detectors (permanence / ballistic / contact) · negative controls include 2 mm tracking noise.

Cohort n Flag rate
Clean negatives 100 0.00
Clean + noise 100 0.00
Teleport / hover / impulse / penetration ~100 each 1.00

Headline: recall 1.000 · precision 1.000 on this calibration set · ~5 s on an M2 laptop for the seeded report.


Pillar 2 — CogVideoX generative failure

Setup: Physics-IQ ball-and-block-fall · CoTracker3 · VIDEO_CHECKS including static_hover · THUDM/CogVideoX-5b-I2V · n=6 (left/center/right × seeds 0–1).

Cohort n Flag rate Checks fired
Real (negative control) 1 0.000
CogVideoX 6 1.000 static_hover × 6

Documented failure (discovery cohort): freeze/hover — tracked object stays airborne with near-zero motion while the prompt requires a fall. Real footage stays quiet. Not Cosmos. Multi-seed discovery cohort (n=6, seeds 0–1); not held-out confirmation — held-out protocol frozen before generation (seeds 2–4 × scenarios 0001–0003). Public rates in pillar2-generative.json.


Document Purpose
Evidence Live charts from metrics API
Methodology report Publishable write-up
Technical article Public distribution article
Design-partner pipeline Outreach + soft-raise status + demand-signal log (async diligence)
Risks Sim-only and evidence-depth risks