A warehouse in San Leandro, California, currently hosts what might be the most delicate game of Jenga in the robotics industry. The building is operated by Encord, a company that builds data tooling for training AI models. Inside, a robotic trainer named Andrew Ceja — the company calls its trainers "pilots" — carefully removes wooden blocks from a teetering tower. He wears a headset with a camera that tracks his gaze, which is fairly standard for collecting robot training data. But this headset also includes sensors that measure his brain waves as he works.
Encord is one of a growing number of startups betting that the next real constraint on humanoid robots and warehouse automation will be the scarcity of real-world physical training data. The company is building a business not just to manage that data but to manufacture it. The brain-wave headset Ceja wears was built by Zander Labs, a German neuroscience startup. Zander's bet is that measuring brain activity — to deduce mental states like error, intent, and surprise — can create a more useful dataset for training models.
The collaboration between Encord and Zander is currently a trial run. Encord says the goal is to build an initial brain-wave-tagged dataset, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up. Lukas Gehrke, a Zander neuroscientist supervising the work, says the amount of brain activity used at any point during a given task offers clues for model builders trying to figure out when they need to deploy their highest-effort models.
Vineeth Velmurugan, Encord's head of robot learning, calls this the "bleeding edge" of the effort to solve the robotics data bottleneck. Velmurugan is a veteran of OpenAI's robot lab and Berkshire Grey, the warehouse automation firm. He joined Encord to build the company's internal data-creation team.
Encord was founded to help companies building machine-vision applications annotate data and evaluate models. As their customers — Velmurugan says they work with many leading robotics firms but is not authorized to name them — began to apply end-to-end learning to robotic manipulation tasks, executives realized they would have to produce training data themselves, rather than simply manage it.
"The data simply does not exist," Velmurugan said.
The bet that generative AI can do for robots what it's done for chatbots keeps running into this same wall. Self-driving car companies collect physical-world data themselves, but that is hard to scale. Training from video can work, but it lacks the fidelity of real-world data. Velmurugan says it will take a dataset something like five times the size of YouTube's video corpus to break through — a scale that helps explain why data-generation itself has become a business and not just a research problem.
Why it matters for European robot service
European robot service providers, integrators, and operators have a direct stake in this data bottleneck. The continent's manufacturing sector, logistics hubs, and service robotics startups all depend on the same underlying constraint: robots need physical training data to learn manipulation tasks, and that data is expensive to produce.
The economics of this problem are stark. Scraping text off the internet, the way LLM makers built their models by pulling from sources like Stack Overflow and the rest of the web, cost frontier labs next to nothing. Generating physical training data does not. This is the limit of the physical-AI-as-LLM comparison. This kind of data has to be manufactured, not just collected, and that changes the economics of building these models.
For European buyers, this means the cost of robot deployment is not just hardware. The training data behind a robot's manipulation skills represents a significant and ongoing investment. Companies like Encord are positioning themselves as the manufacturing layer for that data, drawing egocentric video from several factories around the globe and using their San Leandro facility to experiment with new modalities, like brain waves, or to collect datasets around specific skills for fine-tuning.
The brain-wave approach is particularly interesting for European robotics firms because it comes from a German startup, Zander Labs. The European robotics ecosystem has strong neuroscience and medical technology research traditions, and this trial represents a potential bridge between those fields and physical AI. If brain-wave-tagged datasets prove useful, European companies may have an early advantage in adopting or contributing to this approach.
But the trial is still at an early stage. Encord says the goal is to build an initial brain-wave-tagged dataset, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up. That evaluation has not been completed, and the company has not disclosed results. European operators should treat this as a promising but unproven technique.
The broader trend — data manufacturing as a business — is more established. Encord does both egocentric video collection and robot-remote-operation data collection. When TechCrunch visited, pilots were using leader-follower rigs: paired robotic arms, one controlled directly by a human operator and one that mimics its movements. These rigs were creating data about tasks like pouring coffee from a pot into mugs (very sloshy) and stacking poker chips.
"Every humanoid company has asked us for these pieces," Velmurugan says.
Storage racks at the facility held cartons of fake flowers in vases, books, plastic vegetables, kitty litter trays and scoops, bags and bundles of wires — the stock in trade for training manipulators for household tasks. At one station, another pilot, Sofia Infante, maneuvered robotic arms to plug and unplug ethernet cables from the back of a server — the kind of work data center operators would love to automate, if only robots could manipulate them with the required precision.
A reporter who took a spin behind the controls was able to see why that is still out of reach: pincers are far less dexterous than human fingers and lack the degrees of freedom humans take for granted in their arms.
Another new data modality that Encord is developing uses a set of sensors strapped to the forearm to detect electrical signals in muscles. Video taken of human hands manipulating objects typically does not capture the entire hand, but Velmurugan hopes to build a 3D depiction of where the hand is at any time based on the arm sensors, creating a more robust understanding for models.
Encord's datasets are annotated with physical descriptions of what each video contains — for example, "right hand tightens bolt" — to aid LLM-based models in understanding what is happening. Velmurugan estimates this kind of dense annotation is worth 100 times as much as "junky ego data" for training specific tasks, and it only costs 20 times more to produce, which is a good trade, on paper.
But "20 times more" is still real money, and that is the catch. For European operators, this cost structure matters. Dense annotation may be a good trade on paper, but it is still a significant expense. The question is whether the performance improvement justifies the cost for specific applications.
What buyers and operators should know
For European buyers and operators considering humanoid or warehouse robots, the source material offers several practical takeaways.
First, the data bottleneck is real and central to robot performance. The source material states that the bet that generative AI can do for robots what it's done for chatbots keeps running into the same wall: physical training data is scarce and expensive. Self-driving car companies collect physical-world data themselves, but that is hard to scale. Training from video can work, but it lacks the fidelity of real-world data. Velmurugan says it will take a dataset something like five times the size of YouTube's video corpus to break through.
Second, data generation is now a business, not just a research problem. Companies like Encord are manufacturing training data, not just managing it. This means robot buyers may increasingly see data costs as a line item in their deployment budgets, separate from hardware and software licensing.
Third, the brain-wave approach is unproven. The source material states that Encord's work with Zander is currently a trial run. The goal is to build an initial brain-wave-tagged dataset, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up. No results have been disclosed. Lukas Gehrke, the Zander neuroscientist, says the amount of brain activity used at any point during a given task offers clues for model builders trying to figure out when they need to deploy their highest-effort models. But this is a hypothesis, not a demonstrated outcome.
Fourth, the source material does not disclose which robotics firms Encord works with. Velmurugan says they work with many leading robotics firms but that he is not authorized to name them. Buyers should be aware that claims about industrywide visibility are made by a company with a commercial interest in that positioning.
Fifth, the economics of dense annotation are a trade-off. Velmurugan estimates dense annotation is worth 100 times as much as "junky ego data" for training specific tasks, and it only costs 20 times more to produce. That sounds like a good trade, on paper. But "20 times more" is still real money. For European operators, the question is whether the performance improvement justifies the cost for their specific applications. The source material does not provide specific pricing or ROI figures.
Sixth, the physical limitations of current robot manipulation are significant. The reporter who tried the leader-follower rig found that pincers are far less dexterous than human fingers and lack the degrees of freedom humans take for granted in their arms. This is a reminder that even with better training data, hardware limitations remain a constraint.
Seventh, the facility's work on household tasks — fake flowers, plastic vegetables, kitty litter trays — suggests that the training data market is targeting both household and industrial applications. The ethernet cable plugging task, which a data center operator would want automated, shows the industrial side. European operators in logistics, data centers, and manufacturing should pay attention to which tasks are being trained and whether those tasks match their own needs.
Eighth, the source material does not disclose any specific performance metrics, benchmarks, or customer results for Encord's data products. It does not disclose pricing, delivery times, or service levels. Buyers should not assume any specific performance or cost characteristics beyond what is stated.
Ninth, the source material notes that Encord draws egocentric data from several factories around the globe. This suggests a global supply chain for training data, which may have implications for European companies concerned about data sovereignty or regulatory compliance. The source material does not specify which countries these factories are in.
Tenth, the source material states that Velmurugan says progress is being made — with Encord's visibility into programs across the industry, he is able to see startups and frontier labs alike figure out what works and what doesn't to improve physical AI models. That vantage point — sitting between many robotics companies at once — is also part of Encord's pitch. It can spot which data techniques are gaining traction industrywide before any single customer can. This is a claim about the company's market position, not an independently verified fact.
For European buyers, the practical implication is that the physical AI field is still in a phase where data quality and data cost are the key variables. The brain-wave approach is one of several experimental modalities — alongside forearm muscle sensors and dense annotation — that companies are exploring to improve training data. None of these have been proven at scale yet.
The source material does not provide information about Zander Labs' other customers, the cost of brain-wave headsets, or the timeline for scaling up the trial. It does not disclose whether any European robotics companies are involved in the trial or have expressed interest. It does not provide any regulatory analysis of brain-wave data collection in the European context.
What is known is that a German neuroscience startup and a US-based data tooling company are running a trial to see whether brain-wave-tagged datasets can improve robotics model performance. The trial is at an early stage, and the outcome is not yet known. European operators should monitor this space but should not base purchasing decisions on unproven techniques.
The broader trend — data manufacturing as a business — is more established and more relevant to European buyers. The cost of physical training data is a real constraint on robot deployment, and companies that can manufacture that data efficiently will have a competitive advantage. European operators should factor data costs into their robot deployment plans and should ask their robot vendors how they source and manufacture training data.
The source material also highlights the importance of dense annotation for specific tasks. For European operators deploying robots for narrow, well-defined tasks — like plugging ethernet cables or stacking poker chips — dense annotation may be worth the cost. For broader, less defined tasks, "junky ego data" may be sufficient. The trade-off is a cost-performance decision that each operator must make based on their own requirements.
Finally, the source material does not provide any information about safety, certification, or standards for brain-wave data collection or for the use of such data in robot training. European operators should be aware that this is an emerging area without established standards.
Sources
Published by Vigla Media OÜ (Estonia).