Haga — independent verification for physical AI.
Wedge: Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.
Living prep · July 2026 · Primary capital path is now the bridge track while pending applications and evidence triggers play out. The $1.0M pre-seed pack remains the institutional target in the dataroom; cold outreach on it is paused until an evidence trigger lands.
Problem
Simulated success does not survive deployment. Stanford HAI's 2026 AI Index: robotic manipulation reaches 89.4% success in software sims (RLBench) but only 12% on real household tasks. Physical AI / robotics startups raised $27.6B across 1,009 deals in 2025 — capital is pouring into systems that need verification faster than into verification itself.
Labs and vendors largely self-report sim scores. Synthetic training worlds are rarely checked for physics consistency. Institutional leaderboards (RoboArena) exist at the top of the market, but small teams still lack a fast, commercially accessible, reproducible pre-deployment check.
Solution
An adversarial, physics-grounded evaluation service that stress-tests both artifacts in the physical-AI pipeline:
- Robot policies under randomized mass/friction and multi-tier severity (Pillar 1)
- World-model outputs for physics consistency — permanence, contact, impossible motion (Pillar 2)
Same methodology across both. The product is the private evaluation engagement; methodology, detector definitions, and Lab evidence are the trust signal.
Workflow: policy or generated-world input → adversarial simulation stress-test or physics-consistency pass → pass/fail report with failure cases and thresholds. Integrates into existing sim/release workflows.
Why us / why this lane
Differentiation: the only independent verifier applying the same adversarial method to both generative world-model outputs and robot policies — spanning a pipeline no single competitor currently covers end-to-end.
Open lane: sim-first, reproducible on commodity hardware, narrow-question, fast-turnaround — complementary to Robocurve / RoboArena (real-hardware, institutional scale), not a substitute.
Expansion thesis: Private eval → continuous scoring API → comparative dataset moat → compliance / insurability adjacent. Bottoms-up verification SAM is the wedge; physical-AI capital stack is the ocean — pitch both without inflating TAM arithmetic.
Traction (current evidence)
| Pillar | Snapshot |
|---|---|
| 1 — Policy | Lift 1.00 → 0.26 · Stack 0.96 → 0.20 · PickPlaceCan 1.00 → 0.24 · Door 0.60 → 0.52 (n=50×4, Wilson CIs; mild = gate; severe failures documented) |
| 2 — World model | Checker recall 1.000, FPR 0.000; CogVideoX I2V cohort (n=6, seeds 0–1) flag 1.000 via static_hover vs real 0.000 |
Live charts: Evidence. Written snapshot: Evidence snapshots. Methods: Methodology.
Product & GTM
- Public Lab demonstrates rigor → inbound → scoped private evaluations
- ICP: policy teams needing pre-deployment stress reports; world-model / synthetic-data builders needing physics QA; later enterprise compliance narratives
- Market: bottoms-up verification ~$39M / ~$8M / Y3 ~$0.8M; sells into $27.6B physical-AI capital (2025) as a thin paid layer — detail TAM / SAM / SOM
- Pricing: fixed-scope pilots → packages → continuous scoring (Phase 5); no external dollar quotes until design-partner validation
Team
Sole founder: Mushood Hanif — product, evaluation systems, full-stack. First hire / co-founder search (ML systems / robotics eval) on a written scorecard; join via new option-pool grants, not a placeholder multi-founder roster — Bios · Hire / co-founder scorecard.
Ask posture
Ask room / meetings: layered capital path. Active now: bridge-track first checks / residencies / warm intros (Bridge raise deck). Institutional target remains $1.0M pre-seed for later close — see Use of funds.
No cold capital outreach until an evidence trigger lands (live partner report, grant award, substantive reply, or multi-model held-out). Warm replies / program inboxes always get same-day replies.
See Founder equity & IP, Formation checklist, Financials, and Corporate / legal.