Robot Service Map. Vigla Media OÜ
Media & PR

Nvidia unveils new Cosmos world models, infra for robotics and physical uses – TechCrunch

The intersection of artificial intelligence and physical machinery has long been a domain of incremental progress, but the pace of change is now accelerating in ways that are difficult to overstate. At the SIGGRAPH computer graphics conference in 2025-08, Nvidia made a significant move to consolidate its position in this field, unveiling a comprehensive suite of new world AI models, software libraries, and infrastructure tools aimed squarely at robotics developers and the broader physical AI ecosystem. The announcement, which took place during the conference, signals a deliberate strategy by the company to extend its dominance beyond the data center and into the realm of machines that must perceive, reason, and act within the constraints of the real world.

The centerpiece of this release is Cosmos Reason, a 7-billion-parameter vision language model designed specifically for physical AI applications and robots. This model, as described by Nvidia, is engineered to provide a "reasoning" capability, allowing AI systems to not only see and understand their environment but also to make decisions based on that understanding. This is a notable departure from earlier models that were primarily focused on perception or pattern recognition. The introduction of a reasoning model at this scale suggests a shift toward more autonomous and adaptable robotic systems, ones that can handle novel situations without requiring constant human oversight or pre-programmed responses.

But the announcement was not limited to a single model. Nvidia also introduced a series of neural reconstruction libraries, which are designed to address one of the most challenging aspects of robotics development: simulation. The ability to create accurate, high-fidelity simulations of the real world is critical for training robots, as it allows developers to test and refine their systems in a safe, controlled environment before deploying them in physical settings. The new libraries include a rendering technique that enables developers to simulate the real world in 3D using sensor data. This is a significant advancement, as it moves beyond purely synthetic environments created from scratch and instead allows for the reconstruction of real-world spaces based on actual sensor inputs. This capability is being integrated into CARLA, a popular open-source simulator used by many in the autonomous driving and robotics research communities, making this technology more accessible to a wider range of developers.

The event also featured updates to the Omniverse software development kit, Nvidia's platform for 3D simulation and digital twins. While the specifics of these updates were not fully detailed in the initial announcement, their inclusion underscores the company's commitment to building a comprehensive, end-to-end stack for physical AI development. The Omniverse platform is increasingly seen as a central hub for creating and deploying these technologies, and the updates are likely aimed at improving performance, usability, and integration with the new Cosmos models.

The timing of this announcement is also notable. It comes on the heels of the Cosmos family of world AI models that were first announced at the Consumer Electronics Show (CES) in January of the same year. The rapid iteration and expansion of this product line suggest that Nvidia is moving quickly to establish a dominant position in what it sees as a major growth area. The company's research lab, which has been instrumental in developing these technologies, is now focused on making the models faster and more responsive. As one Nvidia researcher noted, while real-time response is critical for video games and simulations, the requirements for robotics are even more demanding, with reaction times needing to be even faster to be useful in physical applications.

Product and availability details

Beyond the headline-grabbing Cosmos Reason model, Nvidia's announcement included a broader expansion of its Cosmos family of world models. Joining the existing batch are two new models: Cosmos Transfer-2 and a distilled version of Cosmos Transfers. Cosmos Transfer-2 is designed to accelerate synthetic data generation from 3D simulation scenes or spatial control inputs. This is a crucial capability, as the creation of high-quality training data is often a bottleneck in AI development. By making it faster and easier to generate synthetic data from simulations, Nvidia is aiming to reduce the time and cost associated with training robots and AI agents. The distilled version of Cosmos Transfers is optimized for speed, offering a more efficient alternative for developers who prioritize rapid inference over the full capabilities of the larger model.

The company also provided more details on the broader ecosystem, revealing a full-stack approach to physical AI. This includes new open foundation models that are designed to allow robots to reason, plan, and adapt across a wide range of tasks and diverse environments. This is a significant philosophical shift from earlier approaches, which often focused on narrow, task-specific bots. The goal, as articulated by Nvidia, is to move toward more general-purpose robots that can handle a variety of situations without needing to be retrained for each new task. These models are being made available on Hugging Face, a popular platform for sharing and distributing AI models, which should facilitate adoption and community development.

The announcement also touched upon the next generation of Nvidia's Isaac GR00T models, specifically the GR00T N1.6. This is a vision language action (VLA) model that is purpose-built for humanoid robots. The model is designed to enable whole-body control for humanoids, a complex challenge that requires coordinating multiple joints and actuators in a stable and efficient manner. Notably, GR00T relies on Cosmos Reason as its "brain," highlighting the interconnected nature of Nvidia's product stack. The VLA model is a critical piece of the puzzle for humanoid robotics, as it allows the robot to take in visual and linguistic information and translate it into physical actions.

In addition to the models themselves, Nvidia has also focused on the developer experience. The company uploaded a collection of step-by-step guides, inference resources, and post-training workflows to GitHub, collectively referred to as the Cosmos Cookbook. This resource is intended to help developers better understand how to use and train Cosmos models for their specific use cases. The Cookbook covers a range of topics, including data curation, synthetic data generation, and model evaluation. This is a practical move that acknowledges the complexity of deploying these models in real-world applications and aims to lower the barrier to entry for developers who may not have deep expertise in this area.

The models are designed to be used for creating synthetic text, image, and video datasets for training robots and AI agents. This is a core function that underpins many of the other capabilities. By providing a robust set of tools for synthetic data generation, Nvidia is positioning itself as a key supplier of the raw materials needed for the next wave of AI development. The company's push into this area is not just about providing hardware; it is about building an entire ecosystem of software, models, and tools that make it easier for developers to create and deploy physical AI applications.

What it means for buyers

For buyers and developers in the robotics and AI industries, this announcement has several implications. First, it signals a continued and accelerating investment in the tools needed to move AI from the cloud into physical machines. The industry is seeing a broader shift as AI models become capable of learning how to think in the physical world, enabled by cheaper sensors, advanced simulation, and AI models that can increasingly generalize across tasks. Nvidia's move into robotics is a reflection of this trend, and the company is clearly aiming to be a primary supplier of the underlying technology.

The availability of open foundation models on Hugging Face is a particularly significant development for buyers. It means that developers can now access state-of-the-art models without having to build them from scratch. This could significantly reduce the time and cost associated with developing robotic systems, making the technology more accessible to a wider range of companies, from startups to large enterprises. The open nature of these models also fosters a community-driven approach to development, where improvements and innovations can be shared and built upon.

The integration of the neural reconstruction libraries into CARLA is another point of interest. CARLA is already a widely used platform for autonomous driving research, and the addition of this rendering technique could make it an even more powerful tool for simulating real-world scenarios. For buyers, this means that they can potentially create more realistic and accurate simulations, which should lead to better-trained and more reliable systems.

However, it is important to note that the announcement, while extensive, does not disclose all details. For example, specific pricing for the various components of the stack has not been provided. The availability of the models on Hugging Face suggests that at least some of them are open-source, but the commercial terms for other parts of the ecosystem, such as the Omniverse SDK updates or access to the full Cosmos suite, remain unclear. Buyers will need to consult Nvidia directly to understand the full commercial implications of these offerings.

The focus on speed is also a key consideration. As noted by Nvidia's research team, the goal is to make these models respond in real time, with reaction times that are even faster than what is required for video games or simulations. For buyers, this is a critical factor. In a physical environment, a robot that cannot react quickly enough is not just inefficient; it can be dangerous. The emphasis on reducing latency is a clear signal that Nvidia understands the unique demands of physical AI and is working to address them.

The announcement also reflects a realistic perspective from Nvidia's research team. Despite the hype around robots, especially humanoids, the team acknowledges the challenges that remain. The technologies announced are seen as the backbone of the Cosmos family, but there is an understanding that significant work is still needed to bring these systems to full maturity. For buyers, this suggests that while the tools are becoming more powerful, they should still expect to invest time and effort in development and testing.

The move into physical AI is also a strategic play for Nvidia's core business. The company is known for its advanced AI GPUs, and the push into robotics and physical AI is a way to create new demand for these products. As AI models become more complex and are deployed in more demanding physical environments, the need for powerful computing hardware will only grow. By building the software ecosystem that runs on its hardware, Nvidia is creating a moat that makes it more difficult for competitors to displace it.

For buyers, this means that they are not just purchasing a product; they are investing in a platform. The integration between the Cosmos models, the Isaac GR00T models, the Omniverse platform, and the underlying hardware is designed to be seamless. This can be a significant advantage, as it reduces the complexity of integrating disparate tools and systems. However, it also means that buyers may become more locked into Nvidia's ecosystem, which is a consideration that should be weighed carefully.

In summary, the announcement at SIGGRAPH represents a major step forward in the development of physical AI. The introduction of Cosmos Reason, the expansion of the Cosmos family, the new neural reconstruction libraries, and the updates to the broader ecosystem all point to a future where robots are more capable, more adaptable, and more intelligent. For buyers, the key takeaway is that the tools are becoming more powerful and more accessible, but the commercial details and the practical challenges of deployment remain to be fully understood. The pace of innovation is rapid, and those who can effectively leverage these new capabilities will be well-positioned to lead in the next wave of AI-driven automation.

Sources

  • https://techcrunch.com/2025/08/11/nvidia-unveils-new-cosmos-world-models-other-infra-for-physical-applications-of-ai/

Published by Vigla Media OÜ (Estonia).