Human Archive leverages India's gig economy for robotics data

๐กDiscover how a new startup is solving the 'data hunger' of robotics by crowdsourcing real-world human physical data.
โก 30-Second TL;DR
What Changed
Startup founded by Berkeley and Stanford researchers focuses on embodied AI data collection.
Why It Matters
This approach could significantly accelerate the development of general-purpose robots by providing the diverse, real-world datasets currently lacking in simulation-heavy training environments.
What To Do Next
If you are building robotics models, evaluate if your training pipeline requires more diverse real-world human-demonstration data to improve generalization.
Key Points
- โขStartup founded by Berkeley and Stanford researchers focuses on embodied AI data collection.
- โขEmploys gig workers in India to wear camera-equipped caps and sensors.
- โขFocuses on acquiring high-quality real-world physical training data for robotics labs.
- โขAddresses the critical bottleneck of data scarcity in training physical AI agents.
๐ง Deep Insight
Web-grounded analysis with 9 cited sources.
๐ Enhanced Key Takeaways
- โขHuman Archive was founded in 2026 by Raj Patel (CEO), Samay Maini (CTO), Shloke Patel, and Rushil Agarwal, and is based in San Francisco, United States.
- โขThe company secured $500K in seed funding on January 1, 2026, with Y Combinator as an institutional investor.
- โขHuman Archive's data collection infrastructure, launched in March 2026, includes custom hardware rigs, internal models for policy benchmarking, QA pipelines, and annotation software, supported by a 25-person operations team and dedicated AWS servers for terabyte-scale data offloading.
- โขThey offer two primary datasets: HA-Multi, a multimodal dataset with vision, stereo depth, tactile gloves, body IMUs, and wrist cameras, and HA-Ego, a mono RGB vision and wrist cameras dataset, both providing annotations like 3D MANO hand reconstructions, 2D tactile force maps, and human pose reconstructions.
- โขThe startup aims to collect up to 8,000 hours of data daily and has established national-level partnerships to expand its contributor network to over 50,000 individuals, gathering data from diverse environments such as homes, restaurants, hotels, retail, and industrial settings.
๐ ๏ธ Technical Deep Dive
- Data Collection Hardware: Utilizes custom hardware rigs, including camera-equipped caps and sensors, tactile gloves, body IMUs, and wrist cameras.
- Data Modalities: Collects multimodal data including vision (mono RGB and stereo depth with IR dot projection), tactile sensing, and inertial measurement unit (IMU) data.
- Datasets: Offers HA-Multi (fully aligned multimodal data) and HA-Ego (mono RGB vision and wrist cameras data).
- Annotations & Metadata: Provides environment and scene descriptions, high-level task descriptions, task-aligned atomic action labels, hand tracking, object segmentation, SLAM (Simultaneous Localization and Mapping) with extrinsic and intrinsic parameters, and 3D pose reconstruction.
- Customer Outputs: Delivers structured outputs and visualizations such as 3D MANO hand reconstructions, 2D tactile force maps, and depth maps per timestamp.
- Processing Pipeline: Employs internal QA, anonymization, and annotation software, along with internal models for policy benchmarking.
- Infrastructure: Uses dedicated servers at an AWS data center for terabyte-scale data offloading.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (9)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: TechCrunch AI โ

