Living methodology for diligence readers. Public claim surface: methodology, detector definitions / published thresholds, Lab metrics, limited demos. Private: customer artifacts, engagement workflows, partner reports, comparative data — IP strategy. Public article: State of Sim Physics Consistency, v1. Live charts: Evidence. Numeric snapshot: Evidence snapshots.
Figures as of 2026-07-17. Prefer published metrics JSON if schema dates diverge. Sim-only limitation applies to all Pillar 1 numbers.
Claim under test
Haga is an independent verification layer for physical AI: the same adversarial methodology applied to (1) robot policies under physics stress and (2) generative world-model video — without requiring trust in a single demo run.
Wedge (locked): Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.
Pillar 1 — Policy stress (MuJoCo / Robosuite)
Question
Does a deterministic OSC_POSE baseline stay physically consistent — measured by success, grasp failure (where defined), and contact stability — when object / panel mass and contact friction are randomized within documented severity tiers?
Setup
| Field | Value |
|---|---|
| Framework | Robosuite on base MuJoCo |
| Robot / controller | Panda · OSC_POSE |
| Tasks scored | Lift, Stack, PickPlaceCan, Door |
| Episodes | 50 per condition · paired seeds 0..49 · Wilson 95% CIs |
| Stress | Mass + friction (Lift cube / Stack cubeA / PickPlaceCan can / Door panel-scaled) |
| Tiers | mild (gate) · moderate · severe |
| Gate (Lift / Stack / PickPlaceCan) | Stress success ≥ 0.80 × nominal and grasp-fail increase ≤ 0.15 |
| Gate (Door) | Success-primary — success ratio only; grasp N/A for rotating handle |
Full ranges and scoring rules are summarized above; live headline rates publish via the metrics channel to Lab / Evidence charts.
Results (public)
| Task | Nominal → severe | Mild gate | Severe failure | Artifact |
|---|---|---|---|---|
| Lift | 1.00 → 0.26 | PASS | Grasp slip 0.74 | pillar1-lift.json |
| Stack | 0.96 → 0.20 | PASS (≈0.85) | Grasp fail 0.80 | pillar1-stack.json |
| PickPlaceCan | 1.00 → 0.24 | PASS | Grasp slip 0.76 | pillar1-pickplacecan.json |
| Door | 0.60 → 0.52 | PASS (success-primary) | Door did not open 0.48 | pillar1-door.json |
All four tasks publish degradation curves and documented per-seed failure cases (mass / sliding μ / mode) in internal reports. The diligence story is the curve + failures, not a single green badge. Sanitized headline rates: published metrics JSON → Lab / Evidence charts.
Pillar 2 — Physics-consistency checker
Calibration (MuJoCo ground truth)
Position-only detectors — permanence, ballistic, contact — validated against labeled violations before pointing at video.
| Cohort | Flag rate |
|---|---|
| Clean / clean+noise negatives | 0.00 |
| Teleport / hover / impulse / penetration | 1.00 |
Recall 1.000 · precision 1.000 on the calibration set. Detector thresholds are frozen for held-out validation — see Pillar 2 video path below.
Video path (Physics-IQ → CogVideoX)
Pipeline: RGB → CoTracker3 → VIDEO_CHECKS (relaxed permanence / ballistic / contact plus static_hover for generative freeze).
| Layer | Finding |
|---|---|
| Real Physics-IQ negative control | Flag rate 0.000 |
| Synthetic teleport / hover proxies | Flag rate 1.000 |
CogVideoX I2V cohort (n=6, seeds 0–1, THUDM/CogVideoX-5b-I2V) |
Flag rate 1.000 — all six clips fire static_hover |
Documented generative failure: tracked object stays airborne with near-zero motion (freeze/hover) on the ball-and-block-fall prompt. Real footage stays quiet under the same profile. Not Cosmos / NIM — open CogVideoX on free T4 only.
static_hover (frozen VIDEO_CHECKS): airborne fraction ≥ 0.8 · max frame speed < 0.05 m/s · median |a_z + g| on fully-airborne windows ≥ 8.0 m/s² · track length ≥ 11 frames. Held-out protocol frozen before generation (seeds 2–4 × scenarios 0001–0003).
Cohort honesty: the n=6 CogVideoX finding is a post-hoc discovery cohort (static_hover was added after inspecting those clips). It is not the held-out validation result. Held-out protocol is frozen before further clip selection/generation.
What this does not claim
- Real-hardware validation (sim-only until partnered)
- Cosmos / Genie / NIM scores
- Exhaustive multi-policy leaderboard (scripted OSC baselines)
- That mild-gate PASS equals deployment readiness — stress tiers are the product
- Broad generative quality verdict from a 6-clip CogVideoX multi-seed discovery cohort
- That Door’s shallow curve matches Lift/Stack grasp-slip drama — different physics, success-primary gate
- That the discovery cohort substitutes for a pre-registered held-out result
How investors should read this
- Methodology + Lab public as trust; engagements private — inspect ranges, thresholds, seeds, and failure modes here and on the public Lab; partner artifacts and comparative data stay private (IP strategy).
- Show failures — severe-tier grasp slip / place failure and CogVideoX
static_hover(discovery) are intentional proof the verifier can say no. - Same publish channel — Lab, landing, and this room consume the same sanitized metrics feed; raw engineering dumps stay outside the room.
- Gate ≠ story — mild PASS is the entry bar; degradation under moderate/severe is what diligence should quote.
- Discovery ≠ held-out — do not treat the n=6 CogVideoX finding as confirmatory held-out evidence until the frozen protocol cohort is scored.
Next: execute frozen Pillar 2 held-out protocol; convert Priority A replies into live private evaluations (Phase 4). Intake + branded report Ready. Technical article live: State of Sim Physics Consistency, v1. Site submit-for-eval CTA + /methodology remain live.