Living narrative of what the public metrics API currently shows. Live charts stay on the Evidence index (same publish channel as the public Lab / landing). Raw episode dumps never appear in this room.
Figures below mirror fixtures / published artifacts as of 2026-07-17. Prefer the charts if schema dates diverge.
Claim boundaries (read first)
| Bound | What we claim | What we do not claim |
|---|---|---|
| Scope | Sim-first MuJoCo / Robosuite + position-only / tracked-video physics checks | Real-hardware validation |
| Policies | Scripted OSC_POSE baselines on Lift, Stack, PickPlaceCan, Door | Exhaustive multi-policy leaderboard |
| Gate vs stress | Mild is the gate; degradation under moderate/severe is the story | That mild PASS alone proves robustness |
| Door gate | Success-primary (grasp N/A for rotating handle) | Same grasp-fail semantics as Lift/Stack |
| Generative | CogVideoX-5b-I2V cohort, n=6, seeds 0–1, static_hover |
Cosmos / Genie / NIM scores or broad generative quality |
| Ecology | Pre-deployment stress check | Substitute for Robocurve / RoboArena |
Progress digest
| Field | Value |
|---|---|
| Phase | 3 — Public methodology |
| Status | Current (article + CTA live) |
| Next milestone | 2–3 live design-partner evaluation reports |
| Last evidence | 2026-07-17 |
Day-to-day GitHub Issues are not mirrored here.
Pillar 1 — dual-task degradation summary
| Task | Nominal → severe | Mild gate | Severe failure mode | Artifact |
|---|---|---|---|---|
| Lift | 1.00 → 0.26 | PASS | Grasp slip 0.74 | pillar1-lift.json |
| Stack | 0.96 → 0.20 | PASS | Grasp fail 0.80 | pillar1-stack.json |
| PickPlaceCan | 1.00 → 0.24 | PASS | Grasp slip 0.76 | pillar1-pickplacecan.json |
| Door | 0.60 → 0.52 | PASS (success-primary) | Door did not open 0.48 | pillar1-door.json |
Setup (all four): 50 episodes per condition · paired seeds 0..49 · Wilson 95% CIs · Panda · OSC_POSE. Headline rates from published metrics JSON (Lab / Evidence charts).
Pillar 1 — Lift policy stress (Panda)
Setup: 50 episodes per condition · paired seeds · Wilson 95% CIs · gate tier = mild.
| Condition | Success | Success CI | Grasp-fail | Peak EE force (mean) | Verdict |
|---|---|---|---|---|---|
| Nominal | 1.00 | [0.93, 1.00] | 0.00 | 7.30 | — |
| Mild | 1.00 | [0.93, 1.00] | 0.00 | 7.82 | PASS |
| Moderate | 0.90 | [0.79, 0.96] | 0.10 | 12.77 | PASS |
| Severe | 0.26 | [0.16, 0.40] | 0.74 | 14.15 | FAIL |
Headline: success degrades 1.00 → 0.26 from nominal/mild to severe; under severe stress the gripper acquires the cube then loses grasp (slip) in 74% of episodes.
Pillar 1 — Stack policy stress (Panda)
Setup: same tier model · 50 episodes per condition · scripted OSC stack baseline · gate = mild.
| Condition | Success | Success CI | Grasp-fail | Peak EE force (mean) | Verdict |
|---|---|---|---|---|---|
| Nominal | 0.96 | [0.87, 0.99] | 0.04 | 7.44 | — |
| Mild | 0.82 | [0.69, 0.90] | 0.18 | 8.35 | PASS |
| Moderate | 0.72 | [0.58, 0.83] | 0.28 | 12.98 | FAIL |
| Severe | 0.20 | [0.11, 0.33] | 0.80 | 13.58 | FAIL |
Headline: success degrades 0.96 → 0.20; severe grasp-fail 0.80.
Pillar 1 — PickPlaceCan policy stress (Panda)
Setup: same tier model · 50 episodes per condition · gate = mild.
| Condition | Success | Success CI | Grasp-fail | Peak EE force (mean) | Verdict |
|---|---|---|---|---|---|
| Nominal | 1.00 | [0.93, 1.00] | 0.00 | 165.86 | — |
| Mild | 1.00 | [0.93, 1.00] | 0.00 | 214.54 | PASS |
| Moderate | 0.88 | [0.76, 0.94] | 0.12 | 160.43 | PASS |
| Severe | 0.24 | [0.14, 0.37] | 0.76 | 46.87 | FAIL |
Headline: success degrades 1.00 → 0.24; severe grasp-slip 0.76.
Pillar 1 — Door policy stress (Panda)
Setup: panel-scaled mass/friction tiers · success-primary gate (grasp N/A for rotating handle) · 50 episodes per condition.
| Condition | Success | Success CI | Grasp-fail | Peak EE force (mean) | Verdict |
|---|---|---|---|---|---|
| Nominal | 0.60 | [0.46, 0.72] | 0.00 (N/A) | 209.02 | — |
| Mild | 0.64 | [0.50, 0.76] | 0.00 (N/A) | 185.57 | PASS |
| Moderate | 0.58 | [0.44, 0.71] | 0.00 (N/A) | 199.09 | PASS |
| Severe | 0.52 | [0.39, 0.65] | 0.00 (N/A) | 191.91 | PASS |
Headline: success 0.60 → 0.52 — shallower curve than grasp tasks; documented failure mode is door did not open. Gate PASS does not erase the ~48% severe fail rate.
How to read Pillar 1: mild-gate PASS is expected — the diligence story is the degradation curve and documented failure modes across tasks.
Pillar 2 — Physics-consistency checker calibration
Setup: MuJoCo ground-truth trajectories with labeled violations · position-only detectors (permanence / ballistic / contact) · negative controls include 2 mm tracking noise.
| Cohort | n | Flag rate |
|---|---|---|
| Clean negatives | 100 | 0.00 |
| Clean + noise | 100 | 0.00 |
| Teleport / hover / impulse / penetration | ~100 each | 1.00 |
Headline: recall 1.000 · precision 1.000 on this calibration set · ~5 s on an M2 laptop for the seeded report.
Pillar 2 — CogVideoX generative failure
Setup: Physics-IQ ball-and-block-fall · CoTracker3 · VIDEO_CHECKS including static_hover · THUDM/CogVideoX-5b-I2V · n=6 (left/center/right × seeds 0–1).
| Cohort | n | Flag rate | Checks fired |
|---|---|---|---|
| Real (negative control) | 1 | 0.000 | — |
| CogVideoX | 6 | 1.000 | static_hover × 6 |
Documented failure (discovery cohort): freeze/hover — tracked object stays airborne with near-zero motion while the prompt requires a fall. Real footage stays quiet. Not Cosmos. Multi-seed discovery cohort (n=6, seeds 0–1); not held-out confirmation — held-out protocol frozen before generation (seeds 2–4 × scenarios 0001–0003). Public rates in pillar2-generative.json.
Related
| Document | Purpose |
|---|---|
| Evidence | Live charts from metrics API |
| Methodology report | Publishable write-up |
| Technical article | Public distribution article |
| Design-partner pipeline | Outreach + soft-raise status + demand-signal log (async diligence) |
| Risks | Sim-only and evidence-depth risks |