At major technology trade shows over the past year, humanoid robots have become a fixture. They walk, navigate, and manipulate objects with a level of dexterity that would have seemed unrealistic just a few years ago. Yet, according to Gunnar Pétur Hauksson, founder of Treble, a company that develops advanced audio and voice technologies, these machines remain consistently underwhelming when it comes to communication.
Hauksson, who originally trained as a biologist before moving into audio technology, argues that the robotics industry has systematically underdeveloped auditory perception and voice interaction. In a recent editorial published by The Robot Report, he makes the case that this gap is not a minor oversight but a central challenge that could determine whether humanoid robots and mobile physical AI are truly adopted by humans.
The core argument is straightforward: robots that cannot communicate naturally will not be accepted. Hauksson points out that the current competitive landscape in robotics rewards visible, measurable progress. Locomotion is a clear signal of advancement. Vision-based perception has a mature and powerful ecosystem behind it, with abundant data, well-established models, and scalable training pipelines. The entire stack, from data collection to simulation, has evolved to support these modalities. Data is abundant, benchmarks are clear, and improvements are easy to demonstrate.
Simulation, in particular, has become a cornerstone of progress. Platforms like NVIDIA Isaac Sim have enabled rapid iteration and large-scale training in ways that were previously impossible. These systems are powerful, well-designed, and aligned with the broader economics of the industry. But they also reveal something important: the environments used to train intelligent machines are overwhelmingly visual. They are, for the most part, silent.
This is not an accident, Hauksson argues. It reflects a set of rational decisions made under real constraints, including compute limitations, engineering bandwidth, and the need to prioritize what is tractable. But it also means that an entire dimension of perception and interaction has been systematically underdeveloped.
Hauksson draws on human evolution to make the gap clearer. Humans have evolved over millennia to allocate significant energy to processing sensory information. Vision dominates this allocation, accounting for a large portion of the brain's sensory workload. Hearing, by comparison, consumes less. However, it still represents the second most significant share, roughly in the range of 15% to 20%, depending on context.
In the brutal calculus of evolution, energy is never wasted. That 15% to 20% allocation is not an accident, but a direct result of natural selection optimizing our species to survive, thrive, and prosper on planet Earth. Hauksson suggests this should be a glaring hint for roboticists. If a biological intelligence needs that much auditory bandwidth just to navigate and survive in the physical world, silicon intelligence will not succeed without it.
Hearing plays a fundamentally different role than vision. It is central to how humans interpret intent, maintain awareness beyond their field of view, and most importantly, communicate. Through sound, humans infer whether something is approaching or moving away, whether a voice is calm or hostile, and whether an environment is safe or unpredictable. It functions as an always-on layer of perception that complements vision in critical ways.
Speech is not simply a sequence of words. It is a complex exchange of timing, rhythm, micro-intonation, and emotional signaling. It is inherently dynamic and remarkably robust. Humans can communicate effectively in environments that are noisy, reverberant, and chaotic, extracting meaning from sound with a level of resilience that current systems still struggle to match.
Hauksson also highlights a key difference between human and robot communication. Humans benefit from shared biology and deeply ingrained social patterns. They compensate for imperfections in one another's communication because they intuitively understand the system they are part of. Robots do not have this advantage. As a result, they are held to a different standard, particularly in the early stages of adoption.
A robot that moves slightly imperfectly can still be perceived as functional. A robot that communicates poorly, for example, one that mishears, responds out of sync, or fails to operate in real-world acoustic conditions, quickly becomes frustrating or even unsettling. The issue is not just technical performance. It is the breakdown of trust.
The reason this has not been solved, Hauksson argues, is not a lack of awareness but a lack of infrastructure. High-quality audio data is difficult to obtain and even harder to scale. Unlike visual data, it cannot simply be scraped and labeled at scale.
Why it matters for European robot service
For the European robot service industry, Hauksson's argument carries particular weight. Europe is home to a growing number of service robotics companies that are deploying machines in real-world environments: hospitals, warehouses, retail spaces, public transportation hubs, and private homes. These are not controlled laboratory settings. They are loud, complex, and chaotic.
The real world includes crowded trade show floors, industrial settings, city streets, and homes filled with noise, movement of sound sources, reverberation, and unpredictability. These are the environments in which humans operate, and they are precisely the environments where current embodied AI audio and voice systems tend to break down, according to Hauksson.
European service robots are often designed to interact with the public. They guide visitors in museums, deliver meals in hospitals, assist in elderly care facilities, and provide information in airports and train stations. In all of these scenarios, voice communication is not a luxury; it is a core function. A robot that cannot hear a question in a noisy cafeteria, or that misinterprets a command in a reverberant hallway, will not be trusted by its users.
The trust factor is especially important in Europe, where public acceptance of robotics and AI is often more cautious than in other regions. European regulators and consumers tend to place a high premium on safety, transparency, and human-centric design. A robot that communicates poorly is not just a technical failure; it is a social failure. It undermines the very trust that is needed for widespread adoption.
Hauksson's point about the evolutionary allocation of sensory resources also has implications for how European robotics companies should prioritize their development efforts. If hearing represents 15% to 20% of the brain's sensory workload in humans, it is reasonable to expect that a comparable investment in auditory perception and voice interaction is needed for robots that are meant to operate in human environments.
The current focus on vision and locomotion is understandable, but it is incomplete. European companies that are building service robots should consider whether they are allocating sufficient engineering resources to audio and voice. The infrastructure for high-quality audio data is underdeveloped, which means there is a first-mover advantage for companies that invest early.
There is also a broader ecosystem question. Europe has strong research institutions and a growing number of startups working on audio AI, speech recognition, and natural language processing. But these efforts are often fragmented. A more coordinated approach, perhaps at the EU level, could help build the data infrastructure and benchmarks that are needed to advance robotic communication.
The source material does not disclose specific European companies or projects, so it is not possible to name names. What is known is that the challenge is systemic. It affects any company that is building robots meant to interact with humans in real-world acoustic conditions.
What buyers and operators should know
For buyers and operators of service robots, Hauksson's argument offers a practical checklist for evaluation. The first question is not whether a robot can walk or see, but whether it can hear and communicate in the environments where it will actually be used.
Buyers should ask about the robot's performance in noisy, reverberant, and unpredictable acoustic conditions. A robot that works well in a quiet showroom may fail completely on a busy factory floor or in a crowded hospital corridor. The source material does not provide specific test results or performance metrics, so buyers should request their own trials in realistic conditions.
Another key consideration is the quality of the audio data used to train the robot's communication systems. Hauksson notes that high-quality audio data is difficult to obtain and even harder to scale. Unlike visual data, it cannot simply be scraped and labeled at scale. This means that robots trained on limited or synthetic audio data may not generalize well to real-world conditions.
Buyers should also consider the robot's ability to handle the full complexity of human speech. Speech is not just a sequence of words. It involves timing, rhythm, micro-intonation, and emotional signaling. A robot that only processes the literal meaning of words may miss important cues about intent and emotion.
The source material does not disclose specific performance benchmarks, response times, or reliability figures for any particular robot. Buyers should therefore be cautious about any claims that are not backed by transparent testing in realistic environments.
Trust is another critical factor. Hauksson argues that a robot that communicates poorly quickly becomes frustrating or even unsettling. This is not just a matter of user experience; it is a matter of adoption. A robot that breaks down trust will not be used, no matter how well it walks or manipulates objects.
Operators should also think about the long-term maintenance and upgrade path for communication systems. The source material does not disclose specific maintenance requirements, spare-part lead times, or software update policies. Buyers should ask vendors directly about these issues and ensure that they are addressed in service-level agreements.
Finally, buyers should consider the broader ecosystem. The source material notes that the current competitive landscape rewards visible, measurable progress in locomotion and vision. This means that some vendors may have underinvested in audio and voice. Buyers should ask vendors about their audio development roadmap and their investment in this area.
The source material does not provide specific advice on procurement or contracting, so buyers should rely on their own due diligence. What is clear is that communication is not a nice-to-have feature. It is a core capability that will determine whether robots are truly adopted in service environments.
Hauksson's conclusion is direct: human-machine interaction will become the defining hurdle for widespread acceptance of robots by humans because it heavily impacts trust, safety, efficiency, and ease of use. For European buyers and operators, this means that communication should be a top priority in any robot procurement decision.
The source material does not disclose any specific products, vendors, or case studies. What is known is that the challenge is real and systemic. Robots that cannot communicate naturally will not be adopted, regardless of their other capabilities.
Published by Vigla Media OÜ (Estonia).
Sources
Why robots that can’t communicate naturally won’t be adopted