Honest diligence list. Mitigations are sequencing choices, not slogans.
1. Sim-only ecological validity
Risk: Real-hardware competitors (Robocurve, RoboArena) can claim higher ecological validity. Robocurve's public positioning argues simulation “only goes so far.”
Mitigation: Treat sim-only as deliberate v0 scoping for cost and speed. State the limitation next to every public number. Sequence as the pre-deployment check before institutional real-robot evaluation — complementary, not substitute. Locked wedge language: not a public leaderboard.
2. Institutional “trusted evaluator” mindshare
Risk: NVIDIA / Stanford / Berkeley RoboArena occupies trusted third-party mindshare at the top of the market. Claiming “nobody does independent benchmarking” fails a five-minute skeptical search.
Mitigation: Differentiate on speed, accessibility, and commercial service tier — not absence of any independent evaluation. Lead with the open lane: narrow-question, commodity hardware, private engagements.
3. Closest startup competitors moving
Risk: Robocurve (policy, real-hardware) and Instance (video physics) are YC-backed and adjacent. Being early does not guarantee a win.
Mitigation: Dual-artifact coverage (policy + world model under one methodology) is the structural wedge neither owns today. Ship sharper evidence faster than platform ambition.
4. Platform / commoditization (NVIDIA)
Risk: NVIDIA could fold narrow sim checks into Isaac / Cosmos tooling for free, commoditizing the wedge.
Mitigation: Monitor. Win on independent-auditor neutrality and dual-artifact comparative data that a platform vendor is poorly positioned to provide about its own stack. No evidence yet of a lightweight commercial service at Haga's tier — but this is a real ceiling risk.
5. Evidence depth for a seed narrative
Risk: Generative cohort is still a CogVideoX multi-seed case study (n=6, seeds 0–1). Door’s degradation curve is shallower than grasp tasks. Larger scenario sweeps and paid Cosmos-class runs are still open.
Mitigation: Lead with four-task Pillar 1 degradation (Lift/Stack/PickPlaceCan/Door) plus documented CogVideoX static_hover. Soften PASS language; always pair gate with stress. Track remaining gaps in Evidence snapshots and Methodology report.
6. Founder hygiene blocks external sharing
Risk: Unsigned equity / IP makes external invites non-credible and creates assignment risk at incorporation.
Mitigation: Gate A — Founder equity & IP — is a hard gate before Atlas or any raise. Equity 90/10, prior-invention schedule, and side letter are signed (Ready 2026-07-17). Auth still enforces staged investor sharing: investor emails cannot sign in until Atlas + entity docs gates flip.
7. Solo bus-factor
Risk: One founder owns product, harness, and GTM; illness or departure stalls the company.
Mitigation: Written hire / co-founder scorecard (ML systems / robotics eval); quiet search in parallel with evidence and demand; option pool (10% FD locked) for new grants — not a fake multi-founder roster. Post-close milestone M4: first hire or co-founder offer within 9 months of wire (Milestone model); use-of-funds reserves team runway.
8. Replication after proof
Risk: Once value is proven, competitors can reimplement a narrow tool.
Mitigation: Trade-secret harness + accumulating comparative evaluation data + (later) pipeline switching costs. See IP strategy.
9. Market sizing honesty
Risk: Top-down world-model / physical-AI figures can overstate near-term willingness to pay for independent audit.
Mitigation: Lead with the bottoms-up model in TAM / SAM / SOM (~$40M / ~$8M / ~$0.8M). Treat ecosystem capital ($27.6B) and report CAGRs as backdrop only — expansion thesis, not TAM arithmetic. Adjacent WTP (e.g. Sim2Real pilots) is directional, not booked ARR.
Decision stance
At each evidence gate: pivot vertical or framing, not the overall method, when signal is zero. Speed and one sharp reproducible result beat platform ambition.