
What Does XDOF Actually Do?
Teleoperation data is information captured when a human remotely controls a robot to perform a physical task, which is then used to train that robot’s AI models. XDOF specializes in generating exactly this kind of data at scale. Rather than building robots itself, the company builds the pipelines, collection tools, and annotation systems that let AI labs and robotics companies train “general-purpose” robots , machines that can handle varied, everyday physical tasks instead of just one repetitive factory motion.
Think of it like this: large language models like ChatGPT could be trained because the entire internet already existed as a giant pile of text data. Physical robots don’t have that luxury. There’s no equivalent internet-sized dataset of “how humans fold laundry” or “how humans flatten a cardboard box” sitting around waiting to be scraped. Someone has to go out and generate that data, task by task, sensor by sensor. That’s the gap XDOF is filling.
Who Founded XDOF?
XDOF was co-founded in 2024 by Philipp Wu (CEO) and Fred Shentu (CTO), both researchers out of UC Berkeley. Wu’s PhD research focused on how robots learn from large-scale datasets, and he ran into the same wall every robotics researcher eventually hits: there simply wasn’t enough real-world data to work with. Together with Shentu, he built GELLO, a low-cost teleoperation system that lets a single human operator remotely steer a robotic arm to generate usable training data. That project produced an influential research paper and later became the technical foundation for XDOF itself.
Why Is XDOF Raising Again So Soon After Its Series A?
XDOF raised a $70 million Series A in June 2026, with backing from Thrive Capital, Andreessen Horowitz, Lux Capital, and Spark Capital. By most accounts, the company wasn’t planning to go back to investors this quickly. But rapid growth changed the calculus , XDOF’s annualized revenue is reportedly approaching $50 million, a figure that caught the attention of VCs who then approached the company about a new round, rather than the other way around.
This pattern , investors chasing a hot company rather than founders pitching for cash , is becoming increasingly common in the AI infrastructure space, where growth curves have compressed dramatically compared to previous tech cycles.
Is the $1.2 billion valuation confirmed? Not entirely. The terms of the deal are still being finalized and could change before anything is signed. It’s also unclear whether the $1.2B figure already accounts for the new capital being raised or reflects the company’s value before the Series B closes.
The “Scale AI for Robotics” Comparison
Scale AI is a data-labeling company that became essential infrastructure during the LLM boom by supplying the human-annotated data needed to train and fine-tune AI models. Investors are now describing XDOF using a similar comparison , sometimes calling it “the Scale AI or Mercor for physical robotics” , because it plays an analogous role for robots that Scale AI and Mercor played for language models.
The comparison makes sense structurally, but the underlying problem is harder. Text data can be scraped from websites, forums, and books. Physical-world robot data has to be generated, usually by a human physically doing or guiding the task, which makes the entire process slower, more expensive, and far more logistically complex.
How Does XDOF Actually Collect Data?
XDOF combines two main collection methods:
- Remote teleoperation , trained operators steer robotic arms from a distance to perform tasks like folding clothes or flattening boxes
- Egocentric human data capture , collectors wear body sensors that record how a human performs the same everyday tasks, without a robot involved at all
- Global hiring pipelines , the company is building out teams of both teleoperators and sensor-wearing “egocentric operators” across multiple countries
- Annotation systems , raw movement and sensor data is labeled and structured so it’s usable for AI model training
- Academic partnerships , XDOF is working with UC Berkeley’s AI Research lab on a project called ABC, which it describes as aiming to be one of the largest high-quality robot training datasets ever assembled
Who Is Buying This Data?
XDOF has said it already works with around 20 customers, including several frontier AI labs , though the company hasn’t publicly named most of them. This customer base is a big part of why investors are moving so fast: robotics and physical AI companies are increasingly aware that data, not just better algorithms, is the real bottleneck standing between today’s narrow robots and tomorrow’s general-purpose machines.
Why can’t AI labs just collect this data themselves? Building teleoperation hardware, recruiting and training human operators, and running annotation pipelines at scale is a specialized, unglamorous, and operationally heavy business. Most AI labs would rather focus their engineering effort on models and outsource the “dirty work” of physical data collection to a company built specifically for it , which is exactly the niche XDOF occupies.
How Does XDOF Compare to Other Players in the Space?
XDOF isn’t alone in chasing the robot-data opportunity. Here’s how it stacks up against a few notable names:
| Company | Core Focus | Data Collection Method | Notable Backers/Valuation |
| XDOF | Robot teleoperation data pipelines | Remote teleoperation + wearable sensors | 8VC (reported), ~$1.2B (in talks) |
| Scale AI | Originally LLM data labeling, now expanding into robotics/physical data | Human annotation workforce | Long-established data-labeling leader |
| Micro1 | Human-data platform expanding beyond LLMs | Distributed human annotator network | Raised at a ~$500M valuation |
| Mecka AI | Real-world data collection for robot training | Physical/teleoperation-based collection | Emerging competitor in the same niche |
The common thread across all of these companies is the same underlying bet: as AI shifts from purely digital tasks toward physical, embodied tasks, the companies that control high-quality real-world training data will be as important as the labs building the models themselves.
What This Means for the Broader AI/Robotics Industry
Physical AI refers to AI systems that operate in and interact with the real, physical world , robots, autonomous machines, and embodied agents , as opposed to purely digital systems like chatbots. XDOF’s rapid valuation jump is a signal that investors believe physical AI is entering the same “infrastructure land-grab” phase that language models went through a few years ago. Just as picking the right cloud provider or data vendor mattered enormously in the early LLM era, picking the right robot-data partner could shape which companies win the race toward genuinely general-purpose robots.
For students and professionals in India tracking the AI space, this is worth watching closely. Data collection, annotation, and teleoperation are creating an entirely new category of jobs and startups that sit between hardware and pure software AI roles , a hybrid skill space that’s likely to grow fast over the next few years.
FAQ: XDOF’s Series B and Robot Training Data
1. What does XDOF do? XDOF builds the data pipelines, teleoperation tools, and annotation systems used to collect real-world data for training general-purpose robots, acting as an outsourced data supply chain for robotics companies and AI labs.
2. How much is XDOF being valued at in its Series B talks? Reports indicate XDOF is in late-stage talks for a Series B round at a valuation of approximately $1.2 billion, though the deal terms are not yet finalized.
3. Who is leading XDOF’s Series B round? The round is reportedly being led by 8VC, though neither XDOF nor 8VC has publicly confirmed the details.kee
4. How much funding did XDOF raise before this round? XDOF raised a $70 million Series A in June 2026, backed by Thrive Capital, Andreessen Horowitz, Lux Capital, and Spark Capital.
5. Who founded XDOF, and where did the idea come from? XDOF was founded in 2024 by UC Berkeley researchers Philipp Wu (CEO) and Fred Shentu (CTO), building on their earlier teleoperation research project, GELLO.
6. Why is robot training data considered such a big bottleneck? Unlike text-based AI models, which could train on the vast amount of text already available on the internet, physical robots have no equivalent real-world dataset , every bit of physical-task data has to be actively generated through teleoperation or human demonstration.
Want to Keep Up With Stories Like This?
The shift from digital AI to physical, embodied AI is one of the biggest stories in tech right now , and it’s moving fast. If you want to go deeper into how robotics, data pipelines, and physical AI infrastructure actually work, check out Kalinga.ai‘s ongoing workshops and resources on emerging AI trends.