Migrated from haga-core. Technical detail for diligence: Methodology report. Open/proprietary boundary: IP strategy.
Haga — Thesis & Definition
Last verified: July 2026. All figures traced to sources.md unless cited inline.
What Haga Is
Haga is an independent, adversarial, physics-grounded verification layer for physical AI — a service that stress-tests both the simulated worlds physical-AI systems are built on and the policies that act inside them, and reports the results numerically rather than as a self-graded pass.
Physical AI has two failure points that are causally linked but almost never independently checked by the same party:
- The world — generated by a world model or built as a physics simulation — that training and evaluation run inside.
- The agent — a robot policy or embodied controller — acting inside that world.
A physically implausible world teaches an agent the wrong physics. An agent that looks robust inside a flawed simulation fails on deployment. Haga exists to independently verify both — the same adversarial methodology, applied to two connected artifacts, because no other party in the market currently spans both under one roof (see Competitor Analysis).
Haga is not a world-model builder, a robot manufacturer, or a simulation platform. It is the neutral third party that answers: does this system behave consistently under physical stress, and where does it break?
What Haga Does
| Capability | Input | Output |
|---|---|---|
| Policy verification | A trained robot policy (or baseline controller) | Numeric report: success rates, failure modes, contact stability, pass/fail against defined thresholds |
| World-model verification | Synthetic video or generated environments from a generative world model | Physics-consistency score: object permanence, contact plausibility, impossible motion detection |
| Continuous scoring (roadmap) | Customer release pipeline integration | Ongoing physical-plausibility signal at each model/policy version — same posture as independent security scanning in software |
Public posture: methodology, detector definitions / published thresholds, reproducible Lab evidence, and limited demos are the trust signal. Private: customer artifacts, engagement workflows, branded partner reports, and accumulated comparative evaluation data. See IP strategy.
Conversion model: public evidence demonstrates rigor → interested teams contact Haga to submit their own policy or world model for private evaluation. The product is the evaluation engagement, branded partner reports, and (later) continuous scoring — not secrecy of the public claim surface.
How Haga Does It
The methodology is adversarial stress-testing under controlled physical perturbation, applied consistently across two artifact types:
1. Policy verification
A robot policy runs inside simulation under randomized physical conditions:
- Object mass and friction variation
- Sensor noise injection (roadmap)
- Multi-task coverage: Lift, Stack, PickPlace, Door (Robosuite standard tasks)
Behavior is scored against defined thresholds for consistency, robustness, and contact stability. The v0 implementation uses MuJoCo + Robosuite with a locked public benchmark spec; investor synthesis: Methodology report.
2. World-model verification
Generative world-model outputs (synthetic video, simulated environments from Cosmos-class systems) are scored against physics-consistency criteria:
- Object permanence across frames
- Contact plausibility (forces, collisions)
- Physically impossible accelerations or motion
NVIDIA Cosmos is open-weight; evaluation batches can run on rented GPU hours without training new models.
3. Reporting standard
Both artifact types produce the same category of deliverable:
- Reproducible numeric report with defined pass/fail boundaries
- Shown failure cases — not only passes
- Episode counts, seed lists, and variance measures (not single-run screenshots)
- Explicit sim-only limitations stated alongside any result
Published elements (methodology, public Lab metrics, and threshold definitions) are the trust signal. Proprietary elements are partner-specific evaluation artifacts, private comparative datasets from engagements, and commercial continuous-scoring integrations.
Why We Need This
The sim-to-real gap is measured, not theoretical
Stanford HAI's 2026 AI Index Report documents the core problem Haga addresses:
- Robotic manipulation in software simulations (RLBench) has reached 89.4% success
- Robots succeed in only 12% of real household tasks
- The gap between predictable lab settings and unpredictable household environments is wide
Source: Stanford HAI, 2026 AI Index Report — Technical Performance; primary PDF: ai_index_report_2026.pdf.
This is not a marginal accuracy issue. It is an order-of-magnitude trust failure between what systems report in controlled evaluation and what they deliver in deployment. Independent verification exists precisely because self-reporting at this scale is no longer credible.
Capital is flowing into systems that need verification faster than into verification itself
| Signal | Figure | Source |
|---|---|---|
| Physical AI / robotics startup funding (2025) | $27.6B across 1,009 deals — more than 2× 2024 | PitchBook via Mean CEO, May 2026 |
| VC-specific robotics investment (2025) | $7.2B, up from $3.1B in 2023 | Medium physics-simulation analysis, Mar 2026 (citing VC data) |
| AI world models market | $5.8B (2025) → $28.6B (2034), 58.2% CAGR | MarketIntelo, AI World Models Market |
| Interactive physics simulators (closest segment) | ~31.5% of world-models market | MarketIntelo, same report |
| Humanoid robot market | >$6B by 2030 → $51B by 2035, 56% CAGR | Yole Group, May 2026 |
| Broader robotics revenue | $750B by 2035 | Roland Berger |
Note: $27.6B/1,009 deals (broad "physical AI") and $7.2B (VC-specific "robotics") are different scopes — both show the same directional acceleration.
Category validation — adjacent players prove the problem space is real
| Company | What they do | What it validates |
|---|---|---|
| Patronus AI | Simulation-based evaluation of AI agents before deployment | Verification/eval layer is a live commercial category |
| Antioch | Simulation tooling for robot builders | Simulation infrastructure for physical AI is actively built and adopted |
| Bifrost AI | Synthetic labeled 3D data for physical AI | Synthetic-data layer for physical AI is commercially active |
| Instance (YC) | Physics-consistency checking for AI video | Physics-consistency scoring is a live category |
| Robocurve (YC) | Independent robotics benchmarking | Independent benchmarking is a live category |
| Sim2Real | SaaS ($499–$2,500/mo) closing sim-to-real training loops | Willingness to pay for sim-to-real tooling in this problem space |
Sources: sources.md; Sim2Real pricing from public pilot materials, June 2026.
The industry cannot avoid simulation — which makes checking simulation mandatory
NVIDIA CEO Jensen Huang stated at CES 2026 that real-world data collection is "slow, costly, and never enough" — the industry's own leadership acknowledges simulation dependence. NVIDIA's Newton physics engine (2026, built into Isaac Lab) is a direct response to simulation-fidelity limitations, with adopters including ETH Zurich, TU Munich, Boston Dynamics, Figure AI, and Franka Robotics.
Newton improves the simulator. It does not independently audit what comes out of it. Better physics engines and independent verification are parallel investments, not substitutes.
Safety and compliance stakes are codifying now
ANSI/A3 R15.06-2025 — the most significant U.S. industrial robot safety revision in over a decade (published October 2025, replacing the 2012 standard) — harmonizes with ISO 10218-1/2:2025 and shifts from "collaborative robot" as a type to collaborative applications requiring nuanced risk assessment, functional safety validation, and cybersecurity planning.
Sources: A3 announcement; ANSI Blog summary.
Enterprise robotic deployments increasingly require documented safety validation for insurance and compliance. Independent verification intersects directly with insurability as physical AI moves from pilots to production — a stronger commercial argument than market size alone.
Patent activity confirms commercial significance
PatSnap IP analytics (April 2026) identifies sim-to-real solutions as one of the fastest-growing patent areas in industrial robotics — the underlying problem is recognized as commercially significant and defensible, not a niche academic concern.
Structural reasoning (not a cited fact)
Every capital-intensive technology that reaches production scale eventually develops an independent verification layer once self-reporting becomes unacceptable — credit ratings, security audits, food safety certification. Physical AI is entering that phase now. The market figures and sim-to-real gap data above are the evidence; this paragraph is the interpretive frame — not a sourced claim.
Market Analysis
Dedicated diligence page: Market analysis. Summary below.
Trend Analysis
1. Capital acceleration into physical AI
Physical AI and robotics startups raised $27.6B across 1,009 deals in 2025 — more than double 2024 (PitchBook / Mean CEO). VC-specific robotics investment shows the same curve: $7.2B in 2025 vs. $3.1B in 2023.
2. World-model market compounding
AI world models: $5.8B (2025) → $28.6B (2034), 58.2% CAGR. Interactive physics simulators — the segment closest to Haga's target environments — represent ~31.5% of that market (MarketIntelo).
3. Adjacent verticals on the same curve
- Humanoid robots: >$6B by 2030 → $51B by 2035 (56% CAGR, Yole Group)
- Broader robotics revenue: $750B by 2035 (Roland Berger)
- Synthetic data: $710M (2026) → $3.67B (2031) per Mordor Intelligence; $791M → $6.9B per Fortune Business Insights — methodology differs; both confirm an established commercial category
4. Institutional response at the top of the market
NVIDIA co-developed RoboArena with Stanford and UC Berkeley — a live leaderboard evaluating generalist robot policies on real-world tasks, with genuine international competition (Spirit AI took the top spot from NVIDIA's Cosmos 3 model in mid-2026). WorldArena / WorldScore benchmarks embodied world models specifically.
These validate the category — trusted third-party evaluation matters — but operate at institutional scale (real robot fleets, frontier generalist models). Haga's lane is fast-turnaround, reproducible, narrow-question evaluation accessible to teams before they reach RoboArena or Robocurve scale. See competitor-analysis.md.
5. Academic benchmark proliferation (2025–2026)
RoboEval, Polaris, RoboDojo, WorldGym, SC3-Eval — each targets a specific evaluation gap. These set the credibility bar but are published research, not commercial continuous-scoring products.
Importance of This
| Problem | Why it matters for Haga |
|---|---|
| 89.4% sim → 12% real gap (Stanford 2026) | Direct, dramatic evidence that self-reported sim performance is unreliable |
| $27.6B flowing into physical AI (2025) | Large addressable ecosystem of teams that will need verification |
| ANSI/A3 R15.06-2025 compliance pressure | Verification becomes compliance-adjacent, not optional |
| NVIDIA Newton + Cosmos ecosystem | More simulation dependence → more need to audit simulation outputs |
| YC-backed Instance + Robocurve | Category is real and moving — window to establish dual-artifact position |
Lead with the Stanford stat when explaining why independent verification matters. It names the exact failure mode Haga exists to catch better than any market-sizing figure.
Competitor Analysis (IP and Moat)
Full competitive map: competitor-analysis.md. Summary:
| Player | What they do | Overlap with Haga | Relationship |
|---|---|---|---|
| RoboArena (NVIDIA/Stanford/Berkeley) | Public leaderboard, generalist policies, real robot hardware | Institutional policy verification | Different tier — large-scale, not fast-turnaround commercial service |
| Robocurve (YC) | Independent real-hardware robotics benchmarks; open-source "Inspect Robots" | Highest direct overlap on policy side | Real-hardware vs. Haga's sim-first approach; complementary sequencing (Haga pre-deployment, Robocurve post-deployment) |
| Instance (YC) | Physics-consistency scoring for AI-generated video | Closest on world-model side | Video generation only; doesn't extend to policy behavior |
| Sim2Real | SaaS ($499–$2,500/mo) capturing real-world failures → simulation training | Adjacent — training loop optimizer | Different mechanism: customer's own pipeline tool, not independent third-party audit |
| Antioch ($8.5M seed) | Simulation tooling for robot builders | Upstream — builds what Haga verifies | Integration partner, not competitor |
| Bifrost AI ($8.56M total) | Synthetic labeled 3D data | Data generation, not evaluation | Complementary |
| Patronus AI | Digital world models for software agent eval | Same verification thesis, different substrate | Adjacent — not a direct competitor |
| Applied Intuition | Enterprise AV/physical-AI simulation and validation | Adjacent incumbent | Enterprise ceiling; unlikely to compete at Haga's initial wedge |
Haga's wedge (locked one line): Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.
Expansion thesis (venture scale): Private eval → continuous scoring API → comparative dataset moat → compliance / insurability adjacent. Bottoms-up verification SAM is the wedge; physical-AI capital stack is the ocean. Pitch both without inflating TAM arithmetic.
Haga's Moat
-
Accumulated comparative evaluation data. Every stress test builds a private, growing dataset of how policies and world models behave under adversarial conditions — the compounding-data moat credit bureaus and rating agencies have. No single customer sees the aggregate view.
-
Open harness, private comparative data. The public Apache-2.0 benchmark builds trust; the compounding moat is private comparative evaluation data from partner engagements (and later continuous scoring integrations). Patenting a testing methodology is a poor fit — reputation + data accumulate faster.
-
Reputation as neutral third party. Trust compounds with track record — same dynamic as independent security auditors and financial auditors.
-
Integration switching costs. Once a customer's release pipeline depends on continuous Haga scoring, replacing that evaluator requires re-baselining — durable switching cost at subscription scale.
-
Dual-artifact coverage. Policy-only players (Robocurve, RoboArena) and world-model-only players (Instance) cannot replicate the full-pipeline view without building the other pillar from scratch.
Mission and Vision
Mission
Make physical AI systems independently verifiable before they reach deployment — so capital, insurance, and human safety do not depend on self-reported simulation scores.
Vision
Technological: Continuous, API-driven, real-time physical-plausibility scoring spanning both world models and policies — the reference signal the industry cites, embedded the way independent security certifications and credit ratings became infrastructure once their industries matured past self-reporting.
Business: Progression from narrow benchmarking engagements → enterprise evaluation contracts → API-based continuous scoring subscriptions — building toward the category position of established physical-AI infrastructure players (Antioch, Applied Intuition scale) as the independent trust layer builders and their customers rely on rather than build in-house.
How We Start
Haga begins with policy verification methodology — multi-task, multi-severity-tier evaluation with a publicly demonstrable body of results — before extending the same adversarial methodology to generative world-model outputs.
This is deliberate sequencing: prove the methodology is rigorous on the artifact that can be validated fastest, then extend scope with credibility already established. Launching both pillars simultaneously with neither fully proven would weaken the trust signal.
Current state (July 2026): Both pillars have running code and real data. Pillar 1: tiered stress on Lift, Stack, PickPlaceCan, Door (n=50×4, Wilson CIs) — Lift 1.00 → 0.26, Stack 0.96 → 0.20, PickPlaceCan 1.00 → 0.24, Door 0.60 → 0.52 (success-primary gate); severe grasp-slip / place / open failures documented. Pillar 2: checker v0 calibrated — recall 1.000, FPR 0.000; CogVideoX I2V discovery cohort (n=6, seeds 0–1) documents static_hover (post-hoc, not held-out). Held-out protocol v1 frozen before generation. Phase 3 public methodology landed. Open/proprietary posture reconciled. Next: execute held-out matrix; design-partner evals.
Long-Term Goal
| Horizon | Technological | Business |
|---|---|---|
| Near (0–18 mo) | Multi-task benchmark suite, world-model scoring v1, public methodology reports | Design-partner evaluations, inbound from technical distribution |
| Mid (18–36 mo) | Continuous scoring API, customer pipeline integration | Enterprise evaluation contracts |
| Long (3–7 yr) | Industry reference benchmarks cited in papers and procurement | Category leader in independent physical-AI verification |
Target Customers
| Segment | Pain | Haga offer |
|---|---|---|
| Robot policy developers (labs, startups) | Self-reported sim benchmarks don't survive deployment | Pre-deployment stress-test report with shown failure modes |
| World-model builders (Cosmos, Genie, Marble-class) | Synthetic output physics plausibility unverified | Physics-consistency scoring on generated environments/video |
| Synthetic-data vendors | Training data may encode impossible physics | Batch QA on generated datasets |
| Enterprise robotics teams | Compliance and insurance require documented validation | Continuous scoring integrated into release pipeline |
What Haga Is Not
- Not a world-model builder or robot manufacturer
- Not a simulation platform (Antioch/Bifrost territory)
- Not a training-loop optimizer (Sim2Real territory)
- Not a replacement for real-hardware benchmarking (Robocurve/RoboArena territory) — Haga is the fast, reproducible pre-check before teams are ready for institutional real-robot evaluation
- Not a free private evaluation service (public methodology and Lab metrics; partner evals, private reports, and continuous scoring are commercial)
Document Map
| Document | Purpose |
|---|---|
| Thesis | This file — definition, market, mission |
| Market analysis | Trends and why verification matters |
| TAM / SAM / SOM | Bottoms-up ~$40M / ~$8M / ~$0.8M |
| Competitors | Full competitive map with positioning |
| Sources | Citation index for all figures |
| Methodology report | Public methodology synthesis |
| Core specs | Proprietary specs access model (NDA) |
| Roadmap digest | Investor-facing phase progress |