Figures traced to Sources. Competitive map: Competitors. Bottoms-up sizing: TAM / SAM / SOM — ~$40M / ~$8M / ~$0.8M.
Wedge: Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.
Expansion: Private eval → continuous scoring API → comparative dataset moat → compliance / insurability adjacent. Verification SAM is the wedge; physical-AI capital is the ocean.
Trend analysis
1. Capital acceleration into physical AI
Physical AI and robotics startups raised $27.6B across 1,009 deals in 2025 — more than double 2024 (PitchBook / Mean CEO). VC-specific robotics investment shows the same curve: $7.2B in 2025 vs. $3.1B in 2023.
2. World-model market compounding
AI world models: $5.8B (2025) → $28.6B (2034), 58.2% CAGR. Interactive physics simulators — the segment closest to Haga's target environments — represent ~31.5% of that market (MarketIntelo).
3. Adjacent verticals on the same curve
- Humanoid robots: >$6B by 2030 → $51B by 2035 (56% CAGR, Yole Group)
- Broader robotics revenue: $750B by 2035 (Roland Berger)
- Synthetic data: $710M (2026) → $3.67B (2031) per Mordor Intelligence; $791M → $6.9B per Fortune Business Insights — methodology differs; both confirm an established commercial category
4. Institutional response at the top of the market
NVIDIA co-developed RoboArena with Stanford and UC Berkeley — a live leaderboard evaluating generalist robot policies on real-world tasks, with genuine international competition (Spirit AI took the top spot from NVIDIA's Cosmos 3 model in mid-2026). WorldArena / WorldScore benchmarks embodied world models specifically.
These validate the category — trusted third-party evaluation matters — but operate at institutional scale. Haga's lane is fast-turnaround, reproducible, narrow-question evaluation accessible to teams before they reach RoboArena or Robocurve scale.
5. Academic benchmark proliferation (2025–2026)
RoboEval, Polaris, RoboDojo, WorldGym, SC3-Eval — each targets a specific evaluation gap. These set the credibility bar but are published research, not commercial continuous-scoring products.
Why this matters for Haga
| Problem | Why it matters |
|---|---|
| 89.4% sim → 12% real gap (Stanford 2026) | Direct evidence that self-reported sim performance is unreliable |
| $27.6B flowing into physical AI (2025) | Large addressable ecosystem of teams that will need verification |
| ANSI/A3 R15.06-2025 compliance pressure | Verification becomes compliance-adjacent, not optional |
| NVIDIA Newton + Cosmos ecosystem | More simulation dependence → more need to audit simulation outputs |
| YC-backed Instance + Robocurve | Category is real and moving — window to establish dual-artifact position |
Lead with the Stanford stat when explaining why independent verification matters. It names the exact failure mode Haga exists to catch better than any market-sizing figure.
Category validation (adjacent players)
| Company | What they do | What it validates |
|---|---|---|
| Patronus AI | Simulation-based evaluation of AI agents before deployment | Verification/eval layer is a live commercial category |
| Antioch | Simulation tooling for robot builders | Simulation infrastructure for physical AI is actively built |
| Bifrost AI | Synthetic labeled 3D data for physical AI | Synthetic-data layer is commercially active |
| Instance (YC) | Physics-consistency checking for AI video | Physics-consistency scoring is a live category |
| Robocurve (YC) | Independent robotics benchmarking | Independent benchmarking is a live category |
| Sim2Real | SaaS closing sim-to-real training loops | Willingness to pay for sim-to-real tooling |
Full positioning and moat: Competitors. Full narrative context: Thesis.