Haga

Data room

External investor sharing

Market

Competitors

Ready

Competitive map migrated from haga-core/docs/competitor-analysis.md.

Migrated from haga-core. Technical detail for diligence: Methodology report. Open/proprietary boundary: IP strategy.

Competitive Analysis

This document maps every real competitor and adjacent player found in Haga's space, organized by how directly they compete and how much resource they have behind them. All claims are sourced; see inline citations. Last verified: July 2026.

Locked wedge: Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.

Headline finding, stated plainly: the "nobody is independently benchmarking robot policies" framing in the README's Problem section is still true at the narrow, fast-turnaround, commercially accessible layer — but it is no longer true at the top of the market. NVIDIA has co-developed a real, actively-contested leaderboard (RoboArena) with Stanford and Berkeley, and it already has international competition on it. This doesn't kill Haga's thesis, but it means the differentiation has to be sharper than "independent benchmarking doesn't exist" — it has to be "independent benchmarking doesn't exist at the speed, accessibility, and commercial service tier Haga targets."


Tier 0 — Institutional / Academic Benchmarks (the real surprise)

These aren't startups, but they occupy the exact positioning ("independent, trusted, third-party robot capability measurement") Haga's Problem statement claims is empty — and they're backed by the most resourced actors in the field.

Initiative Backers What it does Threat level
RoboArena NVIDIA, Stanford, UC Berkeley Live leaderboard evaluating generalist robot policies on real-world tasks (manipulation, navigation, tool use, perception). Already has genuine international competition — a Chinese lab (Spirit AI) took the top spot from NVIDIA's own Cosmos 3 model in mid-2026, a result covered as real news, not a PR stunt. High — this is the institutional, trusted, cross-company benchmark Haga's own README says doesn't exist. It exists. It's free, credible, and backed by a chip giant plus two top CS departments.
WorldArena / WorldScore Multi-lab consortium (referenced alongside RoboArena) Benchmarks embodied world models specifically — evaluates a model's ability to generate/predict physically plausible worlds from prompts. Directly adjacent to "physics-consistency checking," one of Haga's core value props. Medium-high — different target (world-model generation quality, not policy robustness), but same "trusted evaluator" positioning.
RoboEval, Polaris, RoboDojo, WorldGym, SC3-Eval Various university labs (Stanford, with Amazon AGI / DARPA / NSF funding in some cases) A cluster of 2025–2026 academic benchmarks, each targeting a specific gap: real-to-sim translation for scalable eval (Polaris), structured/scalable manipulation eval (RoboEval), unified sim-and-real benchmarking (RoboDojo), world-model-as-environment for policy eval (WorldGym). Medium — mostly free, open, published research rather than products; not something you buy or integrate, but they set the credibility bar and can be cited by a well-informed customer as "why do I need you when this exists."

What this means for Haga, concretely: these are large-scale, real-robot, foundation-model-focused leaderboards, built for comparing frontier generalist policies (the π0.5/RDT2/Cosmos-tier models) against each other. They require real robot fleets and institutional resources to run. They are not fast-turnaround, not commercially accessible as a private evaluation service, and not usable by a small robotics team wanting a same-day sanity check before shipping. That gap — sim-first, reproducible on commodity hardware, narrow-question, fast — is real and still open. But the "no independent group runs trustworthy benchmarking" line in the README's Problem section should be softened; it's more accurate to say no one runs it at this speed and accessibility tier.


Tier 1 — Direct Startup Competitors

Company Funding Positioning Overlap with Haga
Robocurve YC-backed, early Independent, reproducible, real-hardware robotics benchmarks; open-source "Inspect Robots" framework. Explicitly argues sim-only eval "only goes so far." Highest direct overlap. Same mission statement almost verbatim. Differentiates on real-hardware vs. Haga's sim-first approach — a real, stated tension (see README Risks).
Instance YC-backed, early Physics-consistency quality layer for AI-generated video — scores synthetic video against physics violations. High conceptual overlap (physics-consistency checking) but different artifact (video generation, not robot policies/manipulation). Closer to a "sibling" than a head-to-head competitor.
Sim2Real Live SaaS, $499–$2,500/mo Captures real-world deployment failures and feeds them back into simulation training to close the sim-to-real loop. Different mechanism — an optimization tool for a customer's own training pipeline, not an independent third-party audit. Real, priced, paying-customer precedent (pilot stage as of June 2026) for willingness to pay in this exact problem space.

Tier 2 — Adjacent Infrastructure (potential partners as much as competitors)

Company Funding Positioning Relationship to Haga
Antioch $8.5M seed, $60M valuation (TechCrunch, Apr 2026) Simulation tooling for robot builders without in-house sim capacity ("Cursor for physical AI").
Bifrost AI $8M Series A, $8.56M total (PRNewswire, Oct 2024) Synthetic labeled 3D data generation for physical AI.
Drift / GoDrift Antler-backed pre-seed Robotics-native engineering agent that turns natural language into production-grade simulation workspaces across ROS 2, Gazebo, MuJoCo, Isaac Sim. Developer preview + enterprise demo motion.
Patronus AI $50M Series B, $70M total "Digital World Models" for simulation-based evaluation of AI agents (software/web environments). Same verification thesis (eval layer), different substrate (software agents, not physical robots). Adjacent player, not a direct competitor.

Tier 3 — Platform / Incumbent Risk

Player Why it matters
NVIDIA Co-developer of RoboArena, owner of Cosmos (world model) and Isaac (simulation stack). Has both the resources and stated ambition ("physical AI") to fold narrow benchmarking tools into its ecosystem for free, the way cloud providers commoditize adjacent tooling. If NVIDIA decided to ship a lightweight, accessible version of RoboArena-style checks, that would directly threaten Haga's wedge. No evidence yet that they're doing this at the narrow/local tier — but worth monitoring.
Applied Intuition Large, well-established incumbent in autonomous vehicle and physical AI simulation/validation (named in the FactMR world-model-simulators market report as a major player). Enterprise-focused, high-touch — unlikely to compete for Haga's initial small-team wedge, but a real ceiling on how far upmarket Haga could grow before hitting an entrenched competitor.

Positioning Map

Rendering diagram…

Haga's genuinely open lane sits in the bottom-right quadrant: sim-first, reproducible on commodity hardware, answering one narrow question fast — not competing with RoboArena's real-robot generalist leaderboard, not competing with Robocurve's real-hardware benchmarking-as-a-service, and not competing with Antioch/Bifrost's broader simulation/data platforms. It's the smallest, fastest, cheapest tier of the stack, serving teams too early or too small to use any of the above.


What This Changes About the README

  1. Soften the "nobody does this" claim. The Problem section should acknowledge RoboArena/Polaris/RoboEval exist, and reframe Haga's gap as speed/accessibility, not absence of any independent evaluation.
  2. Robocurve is the closest real competitor, not just an adjacent player. Worth being explicit, including in outreach: Haga is not trying to replace Robocurve's real-hardware benchmarking; it's the pre-deployment sanity check a team runs before it's ready for something like Robocurve or RoboArena.
  3. The "credible authority" positioning is harder to win now. With NVIDIA-Stanford-Berkeley already running a trusted, contested leaderboard, Haga's article and v0 results need to lean on speed and accessibility as the credibility argument, not "we're the only independent check" — that claim no longer survives a skeptical reader's five-minute search, which is exactly the kind of gap this document exists to catch.

Sources