Haga sells independent evaluation, not a simulation platform, training loop, or robot stack.
| Layer | What it is |
|---|---|
| Public posture | Methodology, detector definitions / published thresholds, reproducible Lab evidence, limited demos |
| Product | Private evaluation engagements for policies and/or world-model outputs |
| Conversion | Public Lab evidence → inbound → scoped engagement → numeric report |
What customers buy
A reproducible diligence-grade report that answers: does this system behave consistently under physical stress, and where does it break?
Deliverables
- Numeric scores with defined pass/fail thresholds
- Shown failure cases (not only passes)
- Episode counts, seeds, variance / confidence intervals
- Explicit sim-only limitations stated alongside results
- Optional comparison across policy versions or severity tiers
Motions (today → roadmap)
| Motion | Pillar | Availability |
|---|---|---|
| Policy stress evaluation | 1 | Lift + Stack + PickPlaceCan + Door tiered stress live (n=50×4) |
| World-model physics QA | 2 | Checker calibrated; CogVideoX static_hover failure documented |
| Continuous scoring | Both | Phase 5 — release-pipeline integration |
See Evaluation offering.
Sequencing (deliberate)
- Prove policy verification depth (multi-task, multi-tier, failure cases) — artifact that validates fastest
- Extend the same adversarial methodology to generative world-model batches with credibility already established
- Package as design-partner engagements, then continuous scoring
Launching both pillars as equal “done” claims before either is fully proven would weaken the trust signal. Current state: Phase 1 multi-task (four tasks) and CogVideoX multi-seed generative-failure evidence are live; Phase 3 methodology distribution landed; larger scenario sweeps remain open.
Roadmap: Roadmap digest · Specs: Core specs.
What Haga is not
Wedge reminder: Fast, private, sim-first physics verification for robot policies and generative world-model outputs.
- Not a world-model builder or robot manufacturer
- Not a simulation platform (Antioch / Bifrost territory)
- Not a training-loop optimizer (Sim2Real territory)
- Not a public leaderboard (RoboArena) or real-hardware substitute (Robocurve) — Haga is the fast pre-check before institutional real-robot evaluation
- Not free private evaluation (methodology + harness are the trust signal; partner engagements, reports, and comparative data remain commercial)
Trust model
Same pattern as independent security or financial auditors:
- Inspectable: what was tested, how it was scored, what failed (methodology, published thresholds, Lab metrics, Apache harness)
- Private: customer artifacts, engagement workflows, partner reports, accumulated comparative data — IP strategy