Hypotheses only — refine after Phase 4 design-partner conversations. Do not treat as validated segments.
Primary ICPs
1. Robot policy teams (labs / startups)
- Pain: Self-reported sim benchmarks do not survive deployment; institutional leaderboards are too heavy for a same-week sanity check
- Trigger: Pre-hardware trial, pre-RoboArena / Robocurve spend, investor diligence on policy claims
- Offer: Pillar 1 stress report with tiers, CIs, and shown failure modes
2. World-model / synthetic-data builders
- Pain: Generated environments or video encode impossible physics; buyers ask for QA they cannot self-certify neutrally
- Trigger: Dataset release, model drop, enterprise pilot requiring physics plausibility evidence
- Offer: Pillar 2 physics-consistency scoring on trajectories / tracked video
3. Enterprise robotics groups (later)
- Pain: Compliance / insurance documentation (ANSI/A3-adjacent validation narrative)
- Trigger: Production rollout, insurer or customer audit ask
- Offer: Continuous scoring integrated into release pipeline (Phase 5)
Buying committee (typical)
| Role | Care about |
|---|---|
| Technical lead / research engineer | Methodology rigor, failure cases, reproducibility |
| Founder / PM | Turnaround time, cost vs building in-house |
| Compliance / ops (enterprise) | Documented thresholds, audit trail |
Not ICP (yet)
- Teams that only need real-hardware benchmarking at Robocurve / RoboArena scale — complementary later; do not position as replacement
- Simulation-platform buyers (Antioch / Bifrost territory) — partners, not customers
- Teams seeking training-loop optimization (Sim2Real) — different product category
Disqualification signals
- Wants “PASS badge” without failure cases or CIs
- Requires Cosmos-class scores we have not published
- Expects open-source harness as the deliverable
Conversion packaging: Evaluation offering. Named targets + demand-signal log: Design-partner pipeline.