NVIDIA Cosmos 3: models for reasoning and actions in robots' physical world
NVIDIA released Cosmos 3 — a family of artificial intelligence models for physical systems. Developers can use these models to create robotic applications, autonomous vehicles, and smart spaces that understand the real world, predict what will happen next, and generate actions. Cosmos 3 includes world models, reasoning models, and action models.
AI-processed from NVIDIA Developer Blog; edited by Hamidun News
NVIDIA presented a new generation of its Cosmos platform at the NVIDIA Developer Blog — Cosmos 3, a family of models for so-called Physical AI, which should understand the structure of the real world, predict the development of events in it, and generate actions for robots, autonomous vehicles, and "smart" spaces.
What is Physical AI and why does it need such models
Unlike language models that work with text, or computer vision models that simply recognize objects in images, Physical AI must solve a more complex task — understanding the surrounding world dynamically. NVIDIA formulates this as a three-stage process: the system must first figure out what is happening in its physical environment right now, then predict what is likely to happen next, and only after that generate correct action or action plan. It is precisely this "perception → prediction → action" pipeline that the Cosmos platform targets.
Key facts about the platform:
- Developer — NVIDIA
- Generation name — Cosmos 3
- Task category — "world foundation models" for Physical AI
- Target devices — robots, autonomous vehicles, "smart" spaces
- Announcement source — NVIDIA Developer Blog
How Cosmos fits into the robotics ecosystem
Robots and autonomous vehicles need more than just "seeing" their environment to work safely and effectively — they need a model of the world that allows them to predict the consequences of their own actions and the actions of other participants in the environment before physical movement has already taken place. Previously, such models were developed individually for specific hardware and scenarios, which made them expensive and poorly portable between different types of robots. A general-purpose platform like Cosmos, on the other hand, sets a basic layer of "understanding of physics" of the world, on top of which developers of specific robotic systems — from industrial manipulators to service robots — can build their own more specialized models without starting training from scratch each time.
This approach — a general-purpose model as a foundation on which narrow solutions are then fine-tuned — has already proven its effectiveness in the world of language and visual models, and NVIDIA apparently is betting that the same logic will work for Physical AI as well.
What areas the technology will affect
The practical significance of such models extends far beyond robotics alone. Autonomous transport — the second most important category, where the system must accurately predict the behavior of other road participants fractions of a second in advance to make safe decisions in real time. The third direction — "smart" spaces: production facilities, warehouses and other environments equipped with sensors and cameras, where Physical AI can coordinate the work of multiple devices simultaneously based on a common model of what is happening in the room.
The emergence of such a platform from NVIDIA, one of the leading suppliers of computing infrastructure for AI in general, is a signal that the focus of the industry in 2026 is gradually shifting from models working exclusively with text and images to models that must act in the physical world and bear responsibility for the consequences of their decisions.
For developers of specific robotic systems, the practical value of such a platform lies in reducing the time between idea and working prototype. Instead of spending years collecting and labeling their own dataset of robot interactions with the physical environment, a team can start with an already trained "world foundation" model and adapt it to a specific task — for example, to a certain type of manipulator or warehouse logistics scenario. This is the same path that the computer vision and natural language processing industry took several years ago, when the transition from training models from scratch to fine-tuning ready base models accelerated the launch of applied products to market by an order of magnitude — and NVIDIA apparently expects to repeat this effect for robots and autonomous systems.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.