Back to AI
Startups Are Wiring Human Brains Into Robot Training Loops — Here's Why It Could Matter
AI

Startups Are Wiring Human Brains Into Robot Training Loops — Here's Why It Could Matter

Jul 270 views

Key takeaways

  • Encord is trialing EEG headsets from Zander Labs to capture human brain activity during robot training tasks, hoping neural signals reveal intent and error states that cameras miss.
  • Velmurugan estimates useful physical AI training data requires a corpus roughly five times the size of YouTube, making data generation itself a standalone commercial business.
  • Dense natural language annotation of physical task footage delivers 100x the training value of raw video at only 20x the cost, but the economics still fundamentally differ from cheap LLM text scraping.

Inside a warehouse in San Leandro, California, a worker named Andrew Ceja is carefully pulling wooden blocks from a Jenga tower while wearing a headset that does something unusual: in addition to tracking where his eyes are looking, it reads his brain waves in real time. Ceja is a 'pilot' — the in-house term used by Encord, a data tooling company, for the human trainers who generate physical AI training data. His session represents one of the most experimental bets in robotics right now: that neurological signals can teach machines something that cameras and motion sensors alone cannot.

Encord partnered with Zander Labs, a German neuroscience startup, to build and deploy the EEG-equipped headset. Zander's approach is rooted in the idea that brain activity during a task reveals mental states — moments of error recognition, surprise, or heightened intent — that conventional video capture simply misses. Lucas Gehrke, a Zander neuroscientist overseeing the trial, explains that fluctuations in neural effort throughout a task can signal to robotics model builders exactly when their systems need to engage more sophisticated decision-making resources. The Encord-Zander collaboration is still in trial phase, with the companies planning to tag an initial dataset with brain wave metadata and evaluate whether it measurably improves model performance before committing to a broader rollout.

Vineeth Velmurugan, Encord's head of robot learning and a veteran of both OpenAI's robotics division and warehouse automation firm Berkshire Grey, describes the brain wave work as the 'bleeding edge' of solving what he sees as the central bottleneck in physical AI: the absence of sufficient real-world training data. Velmurugan argues that breaking through to truly capable robotic manipulation will require a dataset roughly five times the size of YouTube's entire video library — a staggering figure that reframes data generation from a research task into a full-blown industry. 'The data simply does not exist,' he told TechCrunch, explaining why Encord shifted from managing customer data to manufacturing it from scratch.

The San Leandro facility runs multiple data collection workflows simultaneously. Alongside the brain wave experiments, pilots operate leader-follower robotic arm rigs — where one arm controlled by a human operator is mirrored by a second robotic arm — to capture manipulation tasks like pouring coffee and stacking poker chips. Another novel modality involves forearm-mounted muscle sensors that detect electrical signals in the tendons, allowing Encord to construct a 3D model of hand position even when camera angles can't capture the full picture. All of this footage is paired with dense natural language annotations — phrases like 'right hand tightens bolt' — which Velmurugan estimates deliver one hundred times the training value of raw, unlabeled footage, despite costing only twenty times more to produce.

The economic reality underlying all of this separates physical AI from the large language model revolution in a fundamental way. LLMs were built by scraping text that already existed on the internet at essentially zero marginal cost. Physical training data has no equivalent shortcut — it must be actively manufactured through human labor, precision hardware, and careful annotation. Encord's position working across many robotics firms simultaneously gives it rare cross-industry visibility into which data techniques are gaining traction, a vantage point the company is turning into a competitive advantage. For Ceja and his colleagues, who previously worked at AI data firm Scale, the work is both technically demanding and genuinely novel — a new kind of occupation being invented in real time to feed the machines of the future.

The bigger picture

The brain wave experiment at Encord is easy to dismiss as a curiosity, but it deserves more serious attention as a signal of where the robotics industry is heading. The fundamental problem — that robots need vastly more embodied, physical-world experience than currently exists — has no clean algorithmic solution. Every workaround, from synthetic data to video scraped from the internet, runs into fidelity problems that compound when you try to translate simulated competence into real-world dexterity. Neurological metadata is a genuinely novel attempt to inject a layer of human cognitive context that video alone strips away, capturing not just what a person did but how hard their brain was working to do it. Whether that signal proves trainable at scale is an open question, but the fact that well-capitalized companies are funding this research at all tells you how seriously the bottleneck is being taken.

The competitive implications are significant. Companies like Figure, Physical Intelligence, and 1X are racing to build general-purpose robots, but all of them depend on data pipelines they don't fully control. Encord's model — becoming the neutral data supplier sitting between competitors — mirrors how cloud infrastructure companies established themselves as indispensable across entire industries. If Encord can demonstrate that its multi-modal, densely annotated datasets meaningfully outperform raw egocentric video, it has the potential to set de facto standards for what 'good' robot training data looks like. That kind of influence compounds over time and becomes very hard for individual robotics firms to replicate internally.

The risk, of course, is that this entire approach gets leapfrogged. Improvements in simulation environments, physics engines, and synthetic data generation are advancing quickly, and some researchers believe the real-to-sim gap will close faster than the data manufacturing industry can scale. Readers should watch for whether major robotics labs begin internalizing data generation operations rather than outsourcing them, and whether the brain wave trial produces any published performance benchmarks — because right now, the promise is compelling but the proof is still pending.

LagPing's take

We're covering Encord's brain wave experiment because it represents a rare convergence of neuroscience, robotics, and the practical economics of AI development — topics that individually get plenty of coverage but rarely collide this explicitly in a single story. At LagPing, we think the data layer of physical AI is dramatically underreported relative to the hardware and model architecture discussions that dominate headlines. Most people know that ChatGPT runs on massive text datasets; far fewer understand that humanoid robots face a data problem so severe that companies are literally paying people to play Jenga in warehouses while their brain activity gets recorded. That's a story worth telling clearly and without hype. We also think the workforce dimension here matters: the 'pilots' at Encord are a new category of tech worker that barely existed five years ago, and understanding what they do tells us something real about how AI is reshaping labor in ways that go beyond the usual displacement narrative. This is the kind of ground-level, infrastructure-focused reporting we want LagPing to be known for.

Shop AI & tech on Amazon

As an Amazon Associate, LagPing earns from qualifying purchases. Product links are affiliate links.

You might also like