Last week, Build AI open-sourced Egocentric-10K: 10,000 hours of first-person footage from factory workers, collected via head-mounted cameras across 85 factories and 2,138 workers.
It’s 16.4 TB of data. 1.08 billion frames. The largest egocentric dataset ever released for robotics.
And it’s completely free, hosted on Hugging Face with Apache 2.0 licensing.
Why would a robotics startup give away this much proprietary data?
Most people will say “community goodwill” or “research contribution.” That’s part of it.
But there’s a deeper GTM strategy at play. One that could define how successful deep tech companies build ecosystems in 2025-26.
Build AI isn’t just releasing data. They’re seeding an ecosystem. And if executed well, this single move could position them as the default platform for embodied AI development.
Here’s the playbook, and what other deep tech companies can learn from it..
Egocentric-10K is the largest egocentric dataset. It is the first dataset collected exclusively in real factories
What Build AI Actually Released
Let’s start with what makes Egocentric-10K significant.
The Data:
10,000 hours of first-person video from real factory workers
1.08 billion frames across 192,900 video clips
Collected from 2,138 workers across 85 factories
State-of-the-art in “hand visibility” and “active manipulation density”
Structured in WebDataset format, ready for ML training
Why This Matters for Robotics:
Most robotics datasets are either:
Simulated (not realistic)
Lab-based (not representative of real work)
Small-scale (hundreds of hours, not thousands)
Egocentric-10K is the first large-scale dataset collected exclusively in real factories, showing actual human manipulation tasks in production environments.
For anyone training embodied AI models, whether for humanoid robots, manipulation tasks, or workflow automation, this is gold.
The Barrier It Removes:
Before Egocentric-10K, if you wanted to train a model on factory manipulation tasks, you’d need to:
Get access to factories
Deploy data collection infrastructure
Coordinate with workers
Handle privacy/legal issues
Process and structure terabytes of data
This would take months and significant capital.
Now: `load_dataset(”builddotai/Egocentric-10K”)` and you’re experimenting in minutes.
Build AI just removed the single biggest barrier to entry for their entire vertical.
The Strategic Logic: Why Open Source This?
Open-sourcing 10,000 hours of proprietary factory data seems counterintuitive.
Build AI clearly invested significant resources:
Hardware (custom Gen 1 head-mounted cameras)
Factory partnerships (access to 85 facilities)
Data collection logistics (coordinating 2,138 workers)
Processing infrastructure (16.4 TB storage, WebDataset formatting)
Why give this away?
The Surface Answer: Accelerate Research
Yes, this helps the research community. Academic labs, startups, and hobbyists can now train embodied AI models without building their own data collection infrastructure.
But Build AI isn’t a non-profit. They’re a venture-backed startup building commercial robotics products.
So what’s the real GTM strategy?
The Deeper Answer: Ecosystem Seeding
Build AI is betting that by open-sourcing Egocentric-10K, they will:
1. Become the Default Platform for Embodied AI Development
When researchers and engineers think “egocentric factory data,” they’ll think Build AI first.
This is brand association. Just like:
ImageNet → Computer vision benchmarking
LAION-5B → Image generation training
Common Crawl → Language model pre-training
Egocentric-10K → Factory robotics / embodied AI
The dataset becomes synonymous with the domain.
2. Drive Adoption of Their Hardware/Platform
Look at the metadata structure. Every video is tagged with:
`”factory_id”`: Unique identifier for the factory
`”worker_id”`: Unique identifier for the worker
Device: “Build AI Gen 1”
Developers training on this data will naturally ask: “How was this collected? What hardware did they use? Can I deploy similar data collection?”
This creates a funnel: Dataset users → Hardware/platform customers
It’s the same playbook as:
Oculus releasing VR datasets → Oculus hardware sales
Apple releasing Core ML models → iPhone/Mac ecosystem lock-in
3. Attract Ecosystem Contributors
By hosting on Hugging Face with Apache 2.0 licensing, Build AI is saying: “Build on this. We want derivatives.”
What happens next:
Researchers fine-tune models on Egocentric-10K
They publish papers citing the dataset
Startups build products using models trained on it
These groups need more data → they approach Build AI
Build AI offers: “Want custom data collection? Use our platform.”
The open dataset seeds commercial relationships.
4. De-Risk Their Own Product Development
Here’s the subtle one: By releasing this data publicly, Build AI gets the entire AI research community working on their problem space.
Academics will:
Find novel ways to use the data
Develop new architectures for egocentric learning
Publish techniques for hand tracking, manipulation detection, etc.
Build AI gets free R&D from the global research community.
Then they can commercialize the best techniques in their products.
This is open-source as competitive intelligence.
5. Create Defensibility Through Network Effects
Once Egocentric-10K becomes the standard benchmark for embodied AI:
Papers will reference it
Models will be trained on it
Derivative datasets will extend it
Build AI becomes the center of the ecosystem.
Even if competitors release similar datasets, Build AI already has:
Brand recognition
Community momentum
Derivative projects building on their data
This is a moat disguised as altruism.
Why This Works Better Than Traditional GTM
Compare Build AI’s approach to traditional deep tech GTM:
Traditional Robotics Company Playbook:
Build proprietary technology
Keep everything closed
Hire sales team
Do enterprise outreach
Close deals one-by-one
Timeline: Years to get traction. High CAC. Limited ecosystem.
Build AI’s Playbook:
Build proprietary technology
Open-source a strategic piece (the dataset)
Let the ecosystem come to you
Convert ecosystem participants to customers
Timeline: Months to ecosystem traction. Low CAC. Self-sustaining network effects.
Why This Works for Deep Tech:
Deep tech products are complex. You can’t “demo” embodied AI in 15 minutes.
Developers need to:
Experiment extensively
Understand the technology deeply
Build trust in the approach
Traditional marketing (blog posts, webinars, sales calls) doesn’t build this trust.
But giving developers something valuable—10,000 hours of real factory data—builds trust immediately.
It signals:
“We have real access to factories” (credibility)
“We understand the data problem” (domain expertise)
“We’re willing to share” (developer-friendly culture)
And it creates reciprocity. When Build AI eventually offers commercial products (hardware, platform services, custom data collection), developers who benefited from the open dataset are more likely to buy.
The Timing is Strategic
Build AI released Egocentric-10K now (November 2024) because:
Embodied AI is exploding (Physical Intelligence just raised $400M, Figure AI hit $2.6B valuation)
There’s a data shortage (everyone needs factory/manipulation data)
They can be first (no comparable open dataset exists at this scale)
Being first with a landmark dataset creates lasting brand association.
The Risk They’re Taking:
Competitors could use this data to build competing products.
But Build AI is betting:
The data alone isn’t the moat (it’s the collection infrastructure, hardware, factory partnerships)
Network effects from open-sourcing outweigh the competitive risk
Being the “default platform” is more valuable than keeping data proprietary
This is the same bet Hugging Face made: open-source the models, monetize the infrastructure.
For deep tech in 2025-26, this is increasingly the winning playbook.
How Other Deep Tech Companies Can Apply This
Build AI’s Egocentric-10K release is a case study in “strategic open source” as GTM.
Here’s how other deep tech companies—especially in AI infrastructure, robotics, and specialized models—can apply this playbook:
Step 1: Identify Your Strategic Asset to Open-Source
Not everything should be open. You need to open-source something that:
- Has high value to developers (removes a major barrier)
- Seeds ecosystem adoption (creates network effects)
- Doesn’t give away your core moat (you retain differentiation)
For Build AI, the dataset was perfect because:
High value (hardest part of embodied AI is data)
Seeds ecosystem (developers train models → need more data/hardware)
Not the core moat (Build AI’s moat is data collection infrastructure, not the data itself)
Examples for other verticals:
The key: Open-source the input (data, tools that help developers start). Keep the platform (infrastructure, deployment, production services) proprietary.
Step 2: Make It Absurdly Easy to Use
Build AI didn’t just dump 16.4 TB of raw video files. They:
Structured it in WebDataset format (ML-ready)
Hosted on Hugging Face (one-line load command)
Documented metadata fields clearly
Provided code examples
Licensed it permissively (Apache 2.0)
This removes all friction. Developers can start experimenting in minutes.
Bad open source: “Here’s our data. Figure it out.”
Good open source: “Here’s one line of code to load it. Here’s what you can build.”
Step 3: Build the Contribution Loop
Open-sourcing data is just the start. You need to create a flywheel:
Dataset release → Developers experiment → They build cool things → They share results → More developers discover dataset → More experiments
Build AI should (and likely will):
Feature projects built on Egocentric-10K
Create leaderboards for models trained on it
Host competitions or challenges
Encourage derivative datasets
Each of these amplifies the ecosystem effect.
Step 4: Convert Ecosystem to Revenue
Open-source datasets don’t directly generate revenue. But they create conversion opportunities:
Developers using the dataset → Potential customers for expanded datasets, custom data collection, hardware
Researchers citing the dataset → Credibility signal for enterprise sales
Startups building on the dataset → Partnership/acquisition targets
Academic collaborations → Talent pipeline, research partnerships
Build AI’s eventual monetization might be:
“Want factory data beyond these 85 facilities? License our platform.”
“Need custom egocentric data for your use case? Deploy our Gen 1 hardware.”
“Training models on Egocentric-10K? Use our inference infrastructure.”
The open dataset is the top of the funnel.
What This Means for Deep Tech in 2025-26
Build AI’s Egocentric-10K release is a signal.
We’re seeing a shift in how deep tech companies think about GTM:
Old Playbook (2020-2023):
Build in stealth
Launch with proprietary tech
Closed ecosystem
Sales-driven growth
New Playbook (2024-2026):
Build in public (or semi-public)
Open-source strategic components
Ecosystem-first approach
Community-driven growth → Enterprise conversion
This shift is happening because:
1. Deep tech is getting more complex
Developers need time to experiment. Traditional “free trial” or “demo” doesn’t work for robotics, specialized AI, or research tools.
Open datasets let developers build real projects, gain confidence, and become evangelists.
2. Network effects matter more than ever
In AI infrastructure and robotics, the platform with the most developers, models, and contributions wins.
Open-sourcing strategic assets accelerates network effects.
3. Developer expectations have changed
Modern developers expect to “try before they buy.” They expect open-source foundations.
Companies that stay fully closed lose developer mindshare to those that open-source strategically.
4. Competition is intensifying
There are now dozens of robotics startups, AI infrastructure companies, and specialized model providers.
Differentiation through technology alone is hard. Differentiation through ecosystem is durable.
The Companies That Will Win in 2025-26:
Not necessarily those with the best technology.
But those with the strongest ecosystems.
And those ecosystems will be seeded by strategic open-source moves like Build AI’s Egocentric-10K.
If you’re building in deep tech and haven’t thought about “what should we open-source?”, you’re already behind.
Build AI Defines a New Era of Ecosystem Playbook
Build AI’s decision to open-source 10,000 hours of factory data isn’t altruism.
It’s strategic GTM.
By removing the biggest barrier to embodied AI development, they’re positioning themselves as the default platform for the entire vertical.
Other deep tech companies should pay attention.
The playbook is clear:
Identify a strategic asset to open-source (data, tools, benchmarks)
Make it absurdly easy to use
Build the contribution loop
Convert ecosystem participants to customers
This is how you build defensible ecosystems in deep tech.
And in 2025-26, ecosystems will matter more than technology alone.
I’m researching how deep tech companies can build developer ecosystems strategically.
If you’re working on AI infrastructure, robotics, or specialized models, I’d love to hear your approach: michelle.aetherone@gmail.com or DM me on X @michellelsun.




