Haga

Data room

External investor sharing

Evidence

Methodology report

Ready

Public methodology write-up — what was tested, how scored, what failed.

Living methodology for diligence readers. Public claim surface: methodology, detector definitions / published thresholds, Lab metrics, limited demos. Private: customer artifacts, engagement workflows, partner reports, comparative data — IP strategy. Public article: State of Sim Physics Consistency, v1. Live charts: Evidence. Numeric snapshot: Evidence snapshots.

Figures as of 2026-07-17. Prefer published metrics JSON if schema dates diverge. Sim-only limitation applies to all Pillar 1 numbers.


Claim under test

Haga is an independent verification layer for physical AI: the same adversarial methodology applied to (1) robot policies under physics stress and (2) generative world-model video — without requiring trust in a single demo run.

Wedge (locked): Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.


Pillar 1 — Policy stress (MuJoCo / Robosuite)

Question

Does a deterministic OSC_POSE baseline stay physically consistent — measured by success, grasp failure (where defined), and contact stability — when object / panel mass and contact friction are randomized within documented severity tiers?

Setup

Field Value
Framework Robosuite on base MuJoCo
Robot / controller Panda · OSC_POSE
Tasks scored Lift, Stack, PickPlaceCan, Door
Episodes 50 per condition · paired seeds 0..49 · Wilson 95% CIs
Stress Mass + friction (Lift cube / Stack cubeA / PickPlaceCan can / Door panel-scaled)
Tiers mild (gate) · moderate · severe
Gate (Lift / Stack / PickPlaceCan) Stress success ≥ 0.80 × nominal and grasp-fail increase ≤ 0.15
Gate (Door) Success-primary — success ratio only; grasp N/A for rotating handle

Full ranges and scoring rules are summarized above; live headline rates publish via the metrics channel to Lab / Evidence charts.

Results (public)

Task Nominal → severe Mild gate Severe failure Artifact
Lift 1.00 → 0.26 PASS Grasp slip 0.74 pillar1-lift.json
Stack 0.96 → 0.20 PASS (≈0.85) Grasp fail 0.80 pillar1-stack.json
PickPlaceCan 1.00 → 0.24 PASS Grasp slip 0.76 pillar1-pickplacecan.json
Door 0.60 → 0.52 PASS (success-primary) Door did not open 0.48 pillar1-door.json

All four tasks publish degradation curves and documented per-seed failure cases (mass / sliding μ / mode) in internal reports. The diligence story is the curve + failures, not a single green badge. Sanitized headline rates: published metrics JSON → Lab / Evidence charts.


Pillar 2 — Physics-consistency checker

Calibration (MuJoCo ground truth)

Position-only detectors — permanence, ballistic, contact — validated against labeled violations before pointing at video.

Cohort Flag rate
Clean / clean+noise negatives 0.00
Teleport / hover / impulse / penetration 1.00

Recall 1.000 · precision 1.000 on the calibration set. Detector thresholds are frozen for held-out validation — see Pillar 2 video path below.

Video path (Physics-IQ → CogVideoX)

Pipeline: RGB → CoTracker3 → VIDEO_CHECKS (relaxed permanence / ballistic / contact plus static_hover for generative freeze).

Layer Finding
Real Physics-IQ negative control Flag rate 0.000
Synthetic teleport / hover proxies Flag rate 1.000
CogVideoX I2V cohort (n=6, seeds 0–1, THUDM/CogVideoX-5b-I2V) Flag rate 1.000 — all six clips fire static_hover

Documented generative failure: tracked object stays airborne with near-zero motion (freeze/hover) on the ball-and-block-fall prompt. Real footage stays quiet under the same profile. Not Cosmos / NIM — open CogVideoX on free T4 only.

static_hover (frozen VIDEO_CHECKS): airborne fraction ≥ 0.8 · max frame speed < 0.05 m/s · median |a_z + g| on fully-airborne windows ≥ 8.0 m/s² · track length ≥ 11 frames. Held-out protocol frozen before generation (seeds 2–4 × scenarios 0001–0003).

Cohort honesty: the n=6 CogVideoX finding is a post-hoc discovery cohort (static_hover was added after inspecting those clips). It is not the held-out validation result. Held-out protocol is frozen before further clip selection/generation.


What this does not claim

  • Real-hardware validation (sim-only until partnered)
  • Cosmos / Genie / NIM scores
  • Exhaustive multi-policy leaderboard (scripted OSC baselines)
  • That mild-gate PASS equals deployment readiness — stress tiers are the product
  • Broad generative quality verdict from a 6-clip CogVideoX multi-seed discovery cohort
  • That Door’s shallow curve matches Lift/Stack grasp-slip drama — different physics, success-primary gate
  • That the discovery cohort substitutes for a pre-registered held-out result

How investors should read this

  1. Methodology + Lab public as trust; engagements private — inspect ranges, thresholds, seeds, and failure modes here and on the public Lab; partner artifacts and comparative data stay private (IP strategy).
  2. Show failures — severe-tier grasp slip / place failure and CogVideoX static_hover (discovery) are intentional proof the verifier can say no.
  3. Same publish channel — Lab, landing, and this room consume the same sanitized metrics feed; raw engineering dumps stay outside the room.
  4. Gate ≠ story — mild PASS is the entry bar; degradation under moderate/severe is what diligence should quote.
  5. Discovery ≠ held-out — do not treat the n=6 CogVideoX finding as confirmatory held-out evidence until the frozen protocol cohort is scored.

Next: execute frozen Pillar 2 held-out protocol; convert Priority A replies into live private evaluations (Phase 4). Intake + branded report Ready. Technical article live: State of Sim Physics Consistency, v1. Site submit-for-eval CTA + /methodology remain live.