Cosmos 3: The Omnimodal Breakthrough in Physical AI World Models

Cosmos 3: Omnimodal World Models for Physical AI

2026-01-01
Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson, Carlos Casanova, Ting-Yun Chang, Yan Chang, Yu-Wei Chao, Prithvijit Chattopadhyay, Roshan Chaudhari, Chieh-Yun Chen, Junyu Chen, Ke Chen, Qizhi Chen, Wenkai Chen, Xiaotong Chen, Yu Chen, An-Chieh Cheng, Click Cheng, Xiu Chia, Jeana Choi, Chaeyeon Chung, Wenyan Cong, Yin Cui, Magdalena Dadela, Nalin Dadhich, Wenliang Dai, Joyjit Daw, Alperen Degirmenci, Rodrigo Vieira Del Monte, Robert Denomme, Sameer Dharur, Marco Di Lucca, Ke Ding, Wenhao Ding, Yifan Ding, Yuzhu Dong, Nicole Drumheller, Yilun Du, Aigul Dzhumamuratova, Aleksandr Efitorov, Hamid Eghbalzadeh, Naomi Eigbe, Imad El Hanafi, Hassan Eslami, Benedikt Falk, Jiaojiao Fan, Jim Fan, Amol Fasale, Sergiy Fefilatyev, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Vikram Fugro, Prashant Gaikwad, TJ Galda, Katelyn Gao, Yihuai Gao, Wenhang Ge, Sreyan Ghosh, Arushi Goel, Vivek Goel, Akash Gokul, Rama Govindaraju, Jinwei Gu, Miguel Guerrero, Elfie Guo, Aryaman Gupta, Siddharth Gururani, Hugo Hadfield, Song Han, Ankur Handa, Zekun Hao, Mohammad Harrim, Ali Hassani, Nathan Hayes-Roth, Yufan He, Chris Helvig, Cyrus Hogg, Madison Huang, Michael Huang, Sophia Huang, Yufan Huang, Jacob Huffman, DeLesley Hutchins, Suneel Indupuru, Boris Ivanovic, Arihant Jain, Joel Jang, Ryan Ji, Yanan Jian, Dongfu Jiang, Jingyi Jin, Atharva Joshi, Nikhilesh Joshi, Pranjali Joshi, Jaehun Jung, Weiwei Kang, Scott Kassekert, Jan Kautz, Ashna Khetan, Julia Kiczka, Slawek Kierat, Gwanghyun Kim, Kuno Kim, Sunny Kim, Kezhi Kong, Xin Kong, Zhifeng Kong, Tomasz Kornuta, Egor Krivov, Hui Kuang, Saurav Kumar, Chia-Wen Kuo, George Kurian, Wojciech Kutak, JF Lafleche, Himangshu Lahkar, Omar Laymoun, Jayjun Lee, Sanggil Lee, Gabriele Leone, Boyi Li, Freya Li, Jiajun Li, Jinfeng Li, Ling Li, Pengcheng Li, Shangru Li, Tingle Li, Xiaolong Li, Xuan Li, Zhaoshuo Li, Zhiqi Li, Hao Liang, Maosheng Liao, Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Sifei Liu, Zihan Liu, Hai Loc Lu, Xiangyu Lu, Alice Luo, Ruipu Luo, Wenjie Luo, Jiangran Lyu, Martin Ding Ma, Nic Ma, Qianli Ma, Dawid Majchrowski, Louis Marcoux, Miguel Martin, Qing Miao, Ashkan Mirzaei, Shreyas Misra, Kaichun Mo, Durra Mohsin, Hyejin Moon, Pawel Morkisz, Saeid Motiian, Kirill Motkov, Seungjun Nah, Yashraj Narang, Deepak Narayanan, Thabang Ngazimbi, Julian Ouyang, David Page, Yatian Pang, Sehwi Park, Mahesh Patekar, Mostofa Patwary, Marco Pavone, Trung Pham, Wei Ping, Soha Pouya, Shrimai Prabhumoye, Varun Praveen, Delin Qu, Hesam Rabeti, Morteza Ramezanali, Marilyn Reeb, Xuanchi Ren, Kristen Rumley, Wojciech Rymer, Jun Saito, Yeongho Seol, John Shao, Piyush Shekdar, Tianwei Shen, Humphrey Shi, Min Shi, Stella Shi, Kevin Shih, Mohammad Shoeybi, Mateusz Sieniawski, Shuran Song, Alexander Sotelo, Amir Sotoodeh, Sunil Srinivasa, Vignesh Srinivasakumar, Bartosz Stefaniak, Rahul Heinrich Steiger, Shangkun Sun, Jiaxiang Tang, Shitao Tang, Yangyang Tang, Yue Tang, Tolou Tavakkoli, Kayley Ting, Krzysztof Tomala, Wei-Cheng Tseng, Jibin Varghese, Sergei Vasilev, Thomas Volk, Raju Wagwani, Roger Waleffe, Andrew Z. Wang, Boxiang Wang, Haoxiang Wang, Qiao Wang, Shihao Wang, Shijie Wang, Ting-Chun Wang, Yan Wang, Yu Wang, David Wehr, Fangyin Wei, Xinshuo Weng, Jay Zhangjie Wu, Kedi Wu, Hongchi Xia, Summer Xiao, Tianjun Xiao, Kevin Xie, Daguang Xu, Jiashu Xu, Mengyao Xu, Ruqing Xu, Xingqian Xu, Yao Xu, Dinghao Yang, Dong Yang, Hans Yang, Xiaodong Yang, Xuning Yang, Yichu Yang, Yurong You, Zhiding Yu, Hao Yuan, Simon Yuen, Xiaohui Zeng, Pengcuo Zeren, Cindy Zha, Haotian Zhang, Jenny Zhang, Jing Zhang, Liangkai Zhang, Paris Zhang, Shun Zhang, Xuanmeng Zhang, Zhizheng Zhang, Ann Zhao, Yilin Zhao, Yuliya Zhautouskaya, Charles Zhou, Fengzhe Zhou, Shilin Zhu, Yuke Zhu, Dima Zhylko, Artur Zolkowski
Summary
Problem
Method
Results
Takeaways
Abstract

NVIDIA introduces Cosmos 3, a family of omnimodal world models (Edge, Nano, Super) that unify multimodal understanding and generation within a single Mixture-of-Transformers (MoT) architecture. It processes language, images, video, audio, and action sequences, establishing a new SOTA for Physical AI backbones that function as both vision-language models and world simulators.

TL;DR

NVIDIA's Cosmos 3 is a paradigm shift in embodied AI. By unifying language, vision, audio, and action into a single Mixture-of-Transformers (MoT) architecture, it eliminates the need for fragmented AI pipelines. It serves as a reasoner, a simulator, and a policy executor all at once, setting new benchmarks in Text-to-Image (T2I), Video Generation, and Robotic Control.

Problem & Motivation: The Fragmentation Bottleneck

In the current state of Physical AI, a robot cleaning a table acts like a patchwork of different "brains":

  1. A VLM to locate the dishes.
  2. A VLA or WAM to calculate arm movements.
  3. A World Model to simulate if the bowl will break if moved.

This fragmented architecture is computationally expensive and logically inconsistent. NVIDIA’s insight is that understanding requires reasoning about the future (generation), and generation relies on an internal model of object persistence and physics (understanding). Cosmos 3 unifies these two pillars, treating "Action" as just another modality to be modeled in a shared latent space.

Methodology: The Mixture-of-Transformers (MoT) Core

1. Dual-Tower Layer Architecture

Cosmos 3 doesn't just pile everything into one big heap. It uses a Mixture-of-Transformers approach. Within each transformer layer, there are two specialized pathways:

  • The Reasoner Tower: Processes Autoregressive (AR) tokens (language and ViT-encoded vision) for understanding.
  • The Generator Tower: Processes Diffusion (DM) tokens (VAE-encoded media, audio, action) for synthesis.

Crucially, these two speak to each other. The Generator tokens utilize bidirectional attention over the AR context, allowing the generated video or action to be perfectly grounded in the text prompt or reasoning trace.

Mixture-of-Transformers (MoT) Architecture Figure: The MoT architecture preserves causal integrity for reasoning while allowing full bidirectional context for diffusion-based generation.

2. Physical Temporal Alignment (Absolute Temporal Modulation)

Handling different sensors is a nightmare because they operate at different speeds (e.g., 24 FPS video vs. 15 Hz robot actions). Cosmos 3 solves this with Absolute Temporal Modulation applied to 3D MRoPE (Multimodal Rotary Positional Embeddings). It assigns temporal coordinates based on real-world time rather than discrete token indices, ensuring that a "second" in video tokens precisely aligns with a "second" in audio and action tokens.

3. Unified Action Tokenization

Whether it's a drone's camera motion, a car's steering, or a robot's 7-DoF arm, Cosmos 3 maps them into a unified action interface. It uses relative transforms (SE(3) poses) and grasp states, allowing the model to learn a "Universal Prior" of movement across wildly different embodiments.

Experiments & Results: SOTA Across the Board

NVIDIA tested Cosmos 3 at three scales: Edge (4B), Nano (16B), and Super (64B).

Reasoning vs. Generation

As shown in the table below, Cosmos 3 is not just "good at everything"—it is often the best. It outperformed Gemini 3.1 Pro and Qwen3-VL in specialized driving and robotics reasoning while simultaneously leading the pack in image and video generation.

Benchmark Comparison Table Figure: Performance of Cosmos 3 against top proprietary and open-source models.

Zero-Shot World Simulation

One of the most impressive feats is Action-Conditioned Generation. When given a robot command (e.g., "put the screwdriver on the shelf"), the model doesn't just guess the next frame; it generates a physically consistent video rollout that matches the executed action.

Robot Policy Performance Figure: Cosmos 3 executing complex multi-step instructions on the RoboArena real-world benchmark.

Critical Analysis & Conclusion

The Power of Synthetic Data (SDG)

A core component of Cosmos 3's success is NVIDIA's Open Synthetic Datasets (SDG-PhyxSim, SDG-DriveSim, etc.). While web-scale data is great for general visuals, it lacks the "long-tail" safety cases needed for Physical AI. By training on high-fidelity simulations of crashes, collisions, and warehouse fires, Cosmos 3 learns physical laws (gravity, momentum) that pure real-world data often misses.

Key Takeaways

  • Unified representation is superior to task-specific pipelines for embodied agents.
  • Action as a Modality: Treating control signals the same as pixels or words allows for massive cross-domain transfer.
  • Efficiency: The MoT design allows for caching the Reasoner tower, dramatically speeding up the generation process.

Limitations & Future Work

The "Sim-to-Real" gap in human motion remains a challenge, as synthetic human data sometimes degrades the model's performance on real human metrics. Future work will likely focus on even tighter integration of proprioceptive feedback for closed-loop control at higher frequencies.

Cosmos 3 proves that world models are not just for generating cool videos—they are the foundational backbones for the next generation of intelligent, physical robots.

Find Similar Papers

Try Our Examples

  • Search for recent papers attempting to unify Autoregressive and Diffusion mechanisms within a single Transformer backbone for multimodal tasks.
  • Which research first introduced the 3D Multimodal Rotary Positional Embedding (MRoPE), and how does Cosmos 3's temporal modulation specifically improve it for variable FPS video generation?
  • Examine recent studies that utilize large-scale synthetic datasets (SDG) for improving physical dynamics and common sense reasoning in vision-language-action models.
Contents
Cosmos 3: The Omnimodal Breakthrough in Physical AI World Models
1. TL;DR
2. Problem & Motivation: The Fragmentation Bottleneck
3. Methodology: The Mixture-of-Transformers (MoT) Core
3.1. 1. Dual-Tower Layer Architecture
3.2. 2. Physical Temporal Alignment (Absolute Temporal Modulation)
3.3. 3. Unified Action Tokenization
4. Experiments & Results: SOTA Across the Board
4.1. Reasoning vs. Generation
4.2. Zero-Shot World Simulation
5. Critical Analysis & Conclusion
5.1. The Power of Synthetic Data (SDG)
5.2. Key Takeaways
5.3. Limitations & Future Work