Figure Claims a Robotics Scaling Law

Figure’s 30-home test challenges whether robot learning truly generalizes beyond demos.
30-Second TL;DR
What Changed
Helix 2.5 performed household tasks in 30 homes without home-specific robot data or fine-tuning.
Why It Matters
The test is a meaningful attempt to measure household generalization outside controlled demos, but a 56% success rate is not yet sufficient for dependable domestic deployment. The debate also shows that robotics scaling claims need standardized, independently verified evaluations.
What To Do Next
Reproduce the evaluation design with held-out homes, objects, and lighting conditions, and report per-task failures rather than only aggregate success.
Key Points
- •Helix 2.5 performed household tasks in 30 homes without home-specific robot data or fine-tuning.
- •The evaluation covered tidying toys, folding towels, and making beds.
- •The reported 56% task success rate triggered criticism from robotics competitors.
Deep Insight
Background and context from public sources — not the original article. 14 sources cited.
Enhanced Key Takeaways
- •Figure AI's Helix 2.5 relies on its crowdsourced 'Index' data ecosystem, which has amassed over 16 million domestic activity videos by paying contributors more than $15 million and ingesting 35 minutes of footage per second.
- •The 56% zero-shot home completion rate represents a 6x performance increase compared to models trained without human video pre-training, which achieved only a 9% success rate.
- •Helix 2.5 cut task-specific fine-tuning data requirements by 50% relative to Helix 02, while no single test task accounted for more than 1.90% of the pre-training distribution.
- •To power its transfer scaling law, Figure entered a $3.5 billion infrastructure agreement with Nscale to secure access to up to 100,000 Nvidia GPUs along the Vera Rubin roadmap.
- •Robotics competitors like Sunday Robotics co-founder Tony Zhao criticized the 44% failure rate as unviable for fragile residential settings, noting their specialized ACT-2 architecture achieves 99.1% success on garment manipulation.
Competitor Analysis
- Primary Training Paradigm
- Human-to-humanoid video pre-training (Index dataset) + sparse adaptation
- Benchmark / Deployment Record
- 56% complete-task success across 30 unmapped homes (237/420 trials)
- Generalization vs. Reliability Trade-off
- High zero-shot environmental breadth, but high (44%) error rate in unstructured consumer settings
- Primary Training Paradigm
- Targeted imitation learning / Action Chunking Transformer
- Benchmark / Deployment Record
- 99.1% success on domestic garment manipulation tasks
- Generalization vs. Reliability Trade-off
- High reliability on constrained manipulation routines, but narrower cross-task zero-shot generalization
| Company / Model | Primary Training Paradigm | Benchmark / Deployment Record | Generalization vs. Reliability Trade-off |
|---|---|---|---|
| Figure AI (Helix 2.5) | Human-to-humanoid video pre-training (Index dataset) + sparse adaptation | 56% complete-task success across 30 unmapped homes (237/420 trials) | High zero-shot environmental breadth, but high (44%) error rate in unstructured consumer settings |
| Sunday Robotics (ACT-2) | Targeted imitation learning / Action Chunking Transformer | 99.1% success on domestic garment manipulation tasks | High reliability on constrained manipulation routines, but narrower cross-task zero-shot generalization |
Technical Deep Dive
- Data Ingestion Engine (Index): Helix 2.5 leverages crowdsourced egocentric and third-person domestic video streams capturing unconstrained real-world behavior, scaling ingestion throughput to approximately 35 minutes of human activity per second.
- Human-to-Humanoid Transfer Scaling Law: The architecture relies on an empirical power-law relationship where action prediction loss scales down predictably with an 8-fold increase in pre-training data, while keeping downstream adaptation capacity fixed.
- Pre-Training Distribution Guardrails: Zero-shot generalization was validated by restricting any single evaluation task (e.g., bed making, towel folding) to under 1.90% of the overall pre-training corpus distribution.
- Compute & Training Infrastructure: Backed by an Nscale multi-year cloud contract provisioned for up to 100,000 Nvidia GPUs (targeting the Vera Rubin generation) to handle continuous multi-modal video tokenization and embodied policy convergence.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2024-02Figure AI closes $675M Series B funding at a $2.6B valuation to accelerate embodied AI development
- 2024-08Figure unveils its second-generation humanoid robot hardware platform, Figure 02
- 2026-08Figure launches the Index crowdsourced video collection ecosystem out of stealth
- 2026-09Figure secures $3.5B compute partnership with Nscale for up to 100,000 Nvidia GPUs
- 2026-09Figure reveals Helix 2.5 and asserts the human-to-humanoid robotics scaling law across 30 test homes
Sources (14)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 极客公园 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.

