Haga

Data room

External investor sharing

Company Overview

Thesis

Ready

Investor-facing thesis narrative migrated from haga-core/docs/thesis.md.

Migrated from haga-core. Technical detail for diligence: Methodology report. Open/proprietary boundary: IP strategy.

Haga — Thesis & Definition

Last verified: July 2026. All figures traced to sources.md unless cited inline.


What Haga Is

Haga is an independent, adversarial, physics-grounded verification layer for physical AI — a service that stress-tests both the simulated worlds physical-AI systems are built on and the policies that act inside them, and reports the results numerically rather than as a self-graded pass.

Physical AI has two failure points that are causally linked but almost never independently checked by the same party:

  1. The world — generated by a world model or built as a physics simulation — that training and evaluation run inside.
  2. The agent — a robot policy or embodied controller — acting inside that world.

A physically implausible world teaches an agent the wrong physics. An agent that looks robust inside a flawed simulation fails on deployment. Haga exists to independently verify both — the same adversarial methodology, applied to two connected artifacts, because no other party in the market currently spans both under one roof (see Competitor Analysis).

Haga is not a world-model builder, a robot manufacturer, or a simulation platform. It is the neutral third party that answers: does this system behave consistently under physical stress, and where does it break?


What Haga Does

Capability Input Output
Policy verification A trained robot policy (or baseline controller) Numeric report: success rates, failure modes, contact stability, pass/fail against defined thresholds
World-model verification Synthetic video or generated environments from a generative world model Physics-consistency score: object permanence, contact plausibility, impossible motion detection
Continuous scoring (roadmap) Customer release pipeline integration Ongoing physical-plausibility signal at each model/policy version — same posture as independent security scanning in software

Public posture: methodology, detector definitions / published thresholds, reproducible Lab evidence, and limited demos are the trust signal. Private: customer artifacts, engagement workflows, branded partner reports, and accumulated comparative evaluation data. See IP strategy.

Conversion model: public evidence demonstrates rigor → interested teams contact Haga to submit their own policy or world model for private evaluation. The product is the evaluation engagement, branded partner reports, and (later) continuous scoring — not secrecy of the public claim surface.


How Haga Does It

The methodology is adversarial stress-testing under controlled physical perturbation, applied consistently across two artifact types:

1. Policy verification

A robot policy runs inside simulation under randomized physical conditions:

  • Object mass and friction variation
  • Sensor noise injection (roadmap)
  • Multi-task coverage: Lift, Stack, PickPlace, Door (Robosuite standard tasks)

Behavior is scored against defined thresholds for consistency, robustness, and contact stability. The v0 implementation uses MuJoCo + Robosuite with a locked public benchmark spec; investor synthesis: Methodology report.

2. World-model verification

Generative world-model outputs (synthetic video, simulated environments from Cosmos-class systems) are scored against physics-consistency criteria:

  • Object permanence across frames
  • Contact plausibility (forces, collisions)
  • Physically impossible accelerations or motion

NVIDIA Cosmos is open-weight; evaluation batches can run on rented GPU hours without training new models.

3. Reporting standard

Both artifact types produce the same category of deliverable:

  • Reproducible numeric report with defined pass/fail boundaries
  • Shown failure cases — not only passes
  • Episode counts, seed lists, and variance measures (not single-run screenshots)
  • Explicit sim-only limitations stated alongside any result

Published elements (methodology, public Lab metrics, and threshold definitions) are the trust signal. Proprietary elements are partner-specific evaluation artifacts, private comparative datasets from engagements, and commercial continuous-scoring integrations.


Why We Need This

The sim-to-real gap is measured, not theoretical

Stanford HAI's 2026 AI Index Report documents the core problem Haga addresses:

  • Robotic manipulation in software simulations (RLBench) has reached 89.4% success
  • Robots succeed in only 12% of real household tasks
  • The gap between predictable lab settings and unpredictable household environments is wide

Source: Stanford HAI, 2026 AI Index Report — Technical Performance; primary PDF: ai_index_report_2026.pdf.

This is not a marginal accuracy issue. It is an order-of-magnitude trust failure between what systems report in controlled evaluation and what they deliver in deployment. Independent verification exists precisely because self-reporting at this scale is no longer credible.

Capital is flowing into systems that need verification faster than into verification itself

Signal Figure Source
Physical AI / robotics startup funding (2025) $27.6B across 1,009 deals — more than 2× 2024 PitchBook via Mean CEO, May 2026
VC-specific robotics investment (2025) $7.2B, up from $3.1B in 2023 Medium physics-simulation analysis, Mar 2026 (citing VC data)
AI world models market $5.8B (2025) → $28.6B (2034), 58.2% CAGR MarketIntelo, AI World Models Market
Interactive physics simulators (closest segment) ~31.5% of world-models market MarketIntelo, same report
Humanoid robot market >$6B by 2030 → $51B by 2035, 56% CAGR Yole Group, May 2026
Broader robotics revenue $750B by 2035 Roland Berger

Note: $27.6B/1,009 deals (broad "physical AI") and $7.2B (VC-specific "robotics") are different scopes — both show the same directional acceleration.

Category validation — adjacent players prove the problem space is real

Company What they do What it validates
Patronus AI Simulation-based evaluation of AI agents before deployment Verification/eval layer is a live commercial category
Antioch Simulation tooling for robot builders Simulation infrastructure for physical AI is actively built and adopted
Bifrost AI Synthetic labeled 3D data for physical AI Synthetic-data layer for physical AI is commercially active
Instance (YC) Physics-consistency checking for AI video Physics-consistency scoring is a live category
Robocurve (YC) Independent robotics benchmarking Independent benchmarking is a live category
Sim2Real SaaS ($499–$2,500/mo) closing sim-to-real training loops Willingness to pay for sim-to-real tooling in this problem space

Sources: sources.md; Sim2Real pricing from public pilot materials, June 2026.

The industry cannot avoid simulation — which makes checking simulation mandatory

NVIDIA CEO Jensen Huang stated at CES 2026 that real-world data collection is "slow, costly, and never enough" — the industry's own leadership acknowledges simulation dependence. NVIDIA's Newton physics engine (2026, built into Isaac Lab) is a direct response to simulation-fidelity limitations, with adopters including ETH Zurich, TU Munich, Boston Dynamics, Figure AI, and Franka Robotics.

Newton improves the simulator. It does not independently audit what comes out of it. Better physics engines and independent verification are parallel investments, not substitutes.

Safety and compliance stakes are codifying now

ANSI/A3 R15.06-2025 — the most significant U.S. industrial robot safety revision in over a decade (published October 2025, replacing the 2012 standard) — harmonizes with ISO 10218-1/2:2025 and shifts from "collaborative robot" as a type to collaborative applications requiring nuanced risk assessment, functional safety validation, and cybersecurity planning.

Sources: A3 announcement; ANSI Blog summary.

Enterprise robotic deployments increasingly require documented safety validation for insurance and compliance. Independent verification intersects directly with insurability as physical AI moves from pilots to production — a stronger commercial argument than market size alone.

Patent activity confirms commercial significance

PatSnap IP analytics (April 2026) identifies sim-to-real solutions as one of the fastest-growing patent areas in industrial robotics — the underlying problem is recognized as commercially significant and defensible, not a niche academic concern.

Structural reasoning (not a cited fact)

Every capital-intensive technology that reaches production scale eventually develops an independent verification layer once self-reporting becomes unacceptable — credit ratings, security audits, food safety certification. Physical AI is entering that phase now. The market figures and sim-to-real gap data above are the evidence; this paragraph is the interpretive frame — not a sourced claim.


Market Analysis

Dedicated diligence page: Market analysis. Summary below.

Trend Analysis

1. Capital acceleration into physical AI

Physical AI and robotics startups raised $27.6B across 1,009 deals in 2025 — more than double 2024 (PitchBook / Mean CEO). VC-specific robotics investment shows the same curve: $7.2B in 2025 vs. $3.1B in 2023.

2. World-model market compounding

AI world models: $5.8B (2025) → $28.6B (2034), 58.2% CAGR. Interactive physics simulators — the segment closest to Haga's target environments — represent ~31.5% of that market (MarketIntelo).

3. Adjacent verticals on the same curve

  • Humanoid robots: >$6B by 2030 → $51B by 2035 (56% CAGR, Yole Group)
  • Broader robotics revenue: $750B by 2035 (Roland Berger)
  • Synthetic data: $710M (2026) → $3.67B (2031) per Mordor Intelligence; $791M → $6.9B per Fortune Business Insights — methodology differs; both confirm an established commercial category

4. Institutional response at the top of the market

NVIDIA co-developed RoboArena with Stanford and UC Berkeley — a live leaderboard evaluating generalist robot policies on real-world tasks, with genuine international competition (Spirit AI took the top spot from NVIDIA's Cosmos 3 model in mid-2026). WorldArena / WorldScore benchmarks embodied world models specifically.

These validate the category — trusted third-party evaluation matters — but operate at institutional scale (real robot fleets, frontier generalist models). Haga's lane is fast-turnaround, reproducible, narrow-question evaluation accessible to teams before they reach RoboArena or Robocurve scale. See competitor-analysis.md.

5. Academic benchmark proliferation (2025–2026)

RoboEval, Polaris, RoboDojo, WorldGym, SC3-Eval — each targets a specific evaluation gap. These set the credibility bar but are published research, not commercial continuous-scoring products.

Importance of This

Problem Why it matters for Haga
89.4% sim → 12% real gap (Stanford 2026) Direct, dramatic evidence that self-reported sim performance is unreliable
$27.6B flowing into physical AI (2025) Large addressable ecosystem of teams that will need verification
ANSI/A3 R15.06-2025 compliance pressure Verification becomes compliance-adjacent, not optional
NVIDIA Newton + Cosmos ecosystem More simulation dependence → more need to audit simulation outputs
YC-backed Instance + Robocurve Category is real and moving — window to establish dual-artifact position

Lead with the Stanford stat when explaining why independent verification matters. It names the exact failure mode Haga exists to catch better than any market-sizing figure.

Competitor Analysis (IP and Moat)

Full competitive map: competitor-analysis.md. Summary:

Player What they do Overlap with Haga Relationship
RoboArena (NVIDIA/Stanford/Berkeley) Public leaderboard, generalist policies, real robot hardware Institutional policy verification Different tier — large-scale, not fast-turnaround commercial service
Robocurve (YC) Independent real-hardware robotics benchmarks; open-source "Inspect Robots" Highest direct overlap on policy side Real-hardware vs. Haga's sim-first approach; complementary sequencing (Haga pre-deployment, Robocurve post-deployment)
Instance (YC) Physics-consistency scoring for AI-generated video Closest on world-model side Video generation only; doesn't extend to policy behavior
Sim2Real SaaS ($499–$2,500/mo) capturing real-world failures → simulation training Adjacent — training loop optimizer Different mechanism: customer's own pipeline tool, not independent third-party audit
Antioch ($8.5M seed) Simulation tooling for robot builders Upstream — builds what Haga verifies Integration partner, not competitor
Bifrost AI ($8.56M total) Synthetic labeled 3D data Data generation, not evaluation Complementary
Patronus AI Digital world models for software agent eval Same verification thesis, different substrate Adjacent — not a direct competitor
Applied Intuition Enterprise AV/physical-AI simulation and validation Adjacent incumbent Enterprise ceiling; unlikely to compete at Haga's initial wedge

Haga's wedge (locked one line): Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.

Expansion thesis (venture scale): Private eval → continuous scoring API → comparative dataset moat → compliance / insurability adjacent. Bottoms-up verification SAM is the wedge; physical-AI capital stack is the ocean. Pitch both without inflating TAM arithmetic.

Haga's Moat

  1. Accumulated comparative evaluation data. Every stress test builds a private, growing dataset of how policies and world models behave under adversarial conditions — the compounding-data moat credit bureaus and rating agencies have. No single customer sees the aggregate view.

  2. Open harness, private comparative data. The public Apache-2.0 benchmark builds trust; the compounding moat is private comparative evaluation data from partner engagements (and later continuous scoring integrations). Patenting a testing methodology is a poor fit — reputation + data accumulate faster.

  3. Reputation as neutral third party. Trust compounds with track record — same dynamic as independent security auditors and financial auditors.

  4. Integration switching costs. Once a customer's release pipeline depends on continuous Haga scoring, replacing that evaluator requires re-baselining — durable switching cost at subscription scale.

  5. Dual-artifact coverage. Policy-only players (Robocurve, RoboArena) and world-model-only players (Instance) cannot replicate the full-pipeline view without building the other pillar from scratch.


Mission and Vision

Mission

Make physical AI systems independently verifiable before they reach deployment — so capital, insurance, and human safety do not depend on self-reported simulation scores.

Vision

Technological: Continuous, API-driven, real-time physical-plausibility scoring spanning both world models and policies — the reference signal the industry cites, embedded the way independent security certifications and credit ratings became infrastructure once their industries matured past self-reporting.

Business: Progression from narrow benchmarking engagements → enterprise evaluation contracts → API-based continuous scoring subscriptions — building toward the category position of established physical-AI infrastructure players (Antioch, Applied Intuition scale) as the independent trust layer builders and their customers rely on rather than build in-house.

How We Start

Haga begins with policy verification methodology — multi-task, multi-severity-tier evaluation with a publicly demonstrable body of results — before extending the same adversarial methodology to generative world-model outputs.

This is deliberate sequencing: prove the methodology is rigorous on the artifact that can be validated fastest, then extend scope with credibility already established. Launching both pillars simultaneously with neither fully proven would weaken the trust signal.

Current state (July 2026): Both pillars have running code and real data. Pillar 1: tiered stress on Lift, Stack, PickPlaceCan, Door (n=50×4, Wilson CIs) — Lift 1.00 → 0.26, Stack 0.96 → 0.20, PickPlaceCan 1.00 → 0.24, Door 0.60 → 0.52 (success-primary gate); severe grasp-slip / place / open failures documented. Pillar 2: checker v0 calibrated — recall 1.000, FPR 0.000; CogVideoX I2V discovery cohort (n=6, seeds 0–1) documents static_hover (post-hoc, not held-out). Held-out protocol v1 frozen before generation. Phase 3 public methodology landed. Open/proprietary posture reconciled. Next: execute held-out matrix; design-partner evals.

Long-Term Goal

Horizon Technological Business
Near (0–18 mo) Multi-task benchmark suite, world-model scoring v1, public methodology reports Design-partner evaluations, inbound from technical distribution
Mid (18–36 mo) Continuous scoring API, customer pipeline integration Enterprise evaluation contracts
Long (3–7 yr) Industry reference benchmarks cited in papers and procurement Category leader in independent physical-AI verification

Target Customers

Segment Pain Haga offer
Robot policy developers (labs, startups) Self-reported sim benchmarks don't survive deployment Pre-deployment stress-test report with shown failure modes
World-model builders (Cosmos, Genie, Marble-class) Synthetic output physics plausibility unverified Physics-consistency scoring on generated environments/video
Synthetic-data vendors Training data may encode impossible physics Batch QA on generated datasets
Enterprise robotics teams Compliance and insurance require documented validation Continuous scoring integrated into release pipeline

What Haga Is Not

  • Not a world-model builder or robot manufacturer
  • Not a simulation platform (Antioch/Bifrost territory)
  • Not a training-loop optimizer (Sim2Real territory)
  • Not a replacement for real-hardware benchmarking (Robocurve/RoboArena territory) — Haga is the fast, reproducible pre-check before teams are ready for institutional real-robot evaluation
  • Not a free private evaluation service (public methodology and Lab metrics; partner evals, private reports, and continuous scoring are commercial)

Document Map

Document Purpose
Thesis This file — definition, market, mission
Market analysis Trends and why verification matters
TAM / SAM / SOM Bottoms-up ~$40M / ~$8M / ~$0.8M
Competitors Full competitive map with positioning
Sources Citation index for all figures
Methodology report Public methodology synthesis
Core specs Proprietary specs access model (NDA)
Roadmap digest Investor-facing phase progress