Skip to content

NVIDIA Cosmos

NVIDIA Cosmos is an open platform of world foundation models and tools for physical AI. Cosmos 3 reasons over images and video and generates video, sound and robot actions; Curator, Evaluator and Cosmos Framework cover data, scoring and post-training. Licenses differ by model (OpenMDW 1.1 for Cosmos 3).1

Also known as Cosmos 3, Cosmos world foundation models (WFMs)

At a glance

What is it?
Cosmos is NVIDIA's family of world foundation models, meaning models trained to understand and predict how the physical world looks and changes, plus open tools around them. Cosmos 3 is a single omnimodal model family built on a Mixture-of-Transformers design. It has two surfaces: a Reasoner that takes text and vision and returns text, and a Generator that takes text, vision, sound and action and produces vision, sound and action. It ships in three base sizes (Super, Nano and Edge). Earlier generations, Cosmos 2 and 2.5, kept perception and generation in separate models.12
What does it do?
Cosmos helps physical AI teams get data and starting models. The Reasoner works as a vision language model for captions, alerts and task planning over video. The Generator produces synthetic video from prompts, images or simulation output, predicts possible futures, and can act as the backbone of a robot policy. Cosmos Curator filters, annotates and deduplicates large sensor datasets, Cosmos Evaluator scores generated outputs, and Cosmos Framework post-trains models with supervised fine-tuning, LoRA, distillation or reinforcement learning.12
Who needs it?
Robotics and autonomous vehicle teams that need more varied training data or a pretrained base for robot policies. Vision AI developers who want a video reasoning model they can run on their own GPUs. Researchers building their own world models, who can reuse the data curation and tokenizer tools.
What does it need?1
  • NVIDIA GPUs sized to the model: H200/B200/GB200 for Super, RTX PRO 6000/H100/B200 for Nano, Jetson AGX Orin/Thor or RTX PRO 6000 for Edge
  • A Python environment (the quickstart uses Python 3.13 with uv) and PyTorch
  • A Hugging Face account with approved access to the gated Cosmos-1.0-Guardrail repository
  • Domain data (video, camera or robot data) for post-training
  • A review process for generated data before it is used for training or safety decisions
What it is not
Cosmos is not a physics simulator with guaranteed accuracy. NVIDIA notes that outputs can show temporal inconsistency, object morphing and implausible dynamics, and that safety-critical uses need extra validation. It is not Omniverse: Omniverse builds 3D simulations, while Cosmos generates and transforms video data and trains physical AI models. It is not only a hosted API, since weights and code are downloadable. Licenses differ by model and distribution: the Cosmos 3 repository releases code and models under OpenMDW 1.1, while NVIDIA's VSS license page lists Cosmos-Reason2 2B and 8B, a Cosmos3-Nano-Reasoner build and Cosmos-Embed1 under the NVIDIA Open Model License, and Cosmos NIM microservices under NVIDIA's software license.123

Availability and licensing. NVIDIA's FAQ states Cosmos world foundation models are available under the OpenMDW 1.1 license from the Linux Foundation, and the Cosmos repository releases source code and models under OpenMDW-1.1, with custom licensing available from NVIDIA. The repository lists Cosmos 3 as released in May 2026 and Cosmos3-Edge in July 2026. Other terms apply to some models: NVIDIA's VSS license page lists Cosmos-Reason2, Cosmos3-Nano-Reasoner and Cosmos-Embed1 under the NVIDIA Open Model License, and Cosmos NIMs under NVIDIA's software license. Check each model.123

The problem it solves

Robots, vehicles and camera systems need training data that covers rare and risky situations: bad weather, unusual objects, near misses. Recording these in the real world is slow, expensive and sometimes unsafe, and each new camera layout or robot body needs fresh data.

Cosmos gives developers pretrained models that already capture how scenes look and move. Teams use them to multiply and vary existing data, to reason over video, and as a starting point for robot policies, then post-train on their own data.

How it works

  1. Curate data. Cosmos Curator processes, annotates, filters and deduplicates video and sensor data.
  2. Pick a base model. Cosmos3-Super (64B), Cosmos3-Nano (16B) or Cosmos3-Edge (4B), depending on quality needs and target hardware.
  3. Reason or generate. The Reasoner answers questions about images and video; the Generator creates video, sound or actions from text, images, video, sound or action input. Simulations from Omniverse can be fed in to produce photoreal variants.
  4. Post-train. Cosmos Framework adapts the model to your cameras, robots or domain; NVIDIA TAO 7 offers agent skills for fine-tuning.
  5. Evaluate. Cosmos Evaluator and benchmark suites score outputs.
  6. Serve. Run through Diffusers, Transformers, vLLM, vLLM-Omni, SGLang, TensorRT LLM or NIM, including OpenAI-compatible endpoints.
NVIDIA Cosmos architecture: components by layer and how they connectApplications &solutionsModels & frameworksInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingReal video and sensor data: Raw input from cameras, vehicles and robotsReal video and sensor dataOmniverse / Isaac Sim simulations: Simulated scenes fed in as control inputOmniverse / Isaac SimsimulationsCosmos 3 Reasoner: Reasons over images and videoCosmos 3 ReasonerCosmos 3 Generator: Generates video, sound and actionsCosmos 3 GeneratorCosmos Framework (post-training): SFT, LoRA, distillation and RL on domain dataCosmos Framework(post-training)Serving: NIM, vLLM, SGLang, TensorRT LLM: Runs models behind APIsServing: NIM, vLLM, SGLang,TensorRT LLMCosmos Curator: Filters, annotates and deduplicates dataCosmos CuratorCosmos Evaluator: Scores generated outputsCosmos EvaluatorNVIDIA GPUs from Jetson to GB200: Hardware tiers for Edge, Nano and SuperNVIDIA GPUs from Jetson toGB200
Diagram as a list
  1. Applications & solutions

    • Real video and sensor dataRaw input from cameras, vehicles and robotsConnects to Cosmos Curator
    • Omniverse / Isaac Sim simulationsSimulated scenes fed in as control inputConnects to Cosmos 3 Generator
  2. Models & frameworks

    • Cosmos 3 ReasonerReasons over images and videoConnects to Serving: NIM, vLLM, SGLang, TensorRT LLM
    • Cosmos 3 GeneratorGenerates video, sound and actionsConnects to Cosmos Evaluator, Serving: NIM, vLLM, SGLang, TensorRT LLM
    • Cosmos Framework (post-training)SFT, LoRA, distillation and RL on domain dataConnects to Cosmos 3 Reasoner, Cosmos 3 Generator
  3. Inference & runtime software

    • Serving: NIM, vLLM, SGLang, TensorRT LLMRuns models behind APIs
  4. Operations & orchestration

    • Cosmos CuratorFilters, annotates and deduplicates dataConnects to Cosmos Framework (post-training)
    • Cosmos EvaluatorScores generated outputs
  5. Accelerated computing

    • NVIDIA GPUs from Jetson to GB200Hardware tiers for Edge, Nano and SuperConnects to Serving: NIM, vLLM, SGLang, TensorRT LLM
Components and connections as documented by NVIDIA.

Capabilities

  • World reasoning (Reasoner)12

    Takes text and images or video and returns text, for world understanding, grounding, task planning and embodied reasoning.

    Why it matters: Gives video analytics agents and robots a model that can describe and reason about scenes.

    Limits: Answers can be wrong; safety-critical uses need extra validation.

  • World generation (Generator)1

    Takes text, vision, sound and action input and generates video, sound and actions for simulation, future prediction and synthetic data.

    Why it matters: Expands training data beyond what was physically recorded.

    Limits: Long, high-resolution or physically complex outputs can contain artifacts; diffusion steps are compute-heavy.

  • Three base model sizes1

    Cosmos3-Super (64B), Cosmos3-Nano (16B) and Cosmos3-Edge (4B) share one architecture across data center, workstation and edge tiers.

    Why it matters: The same model family can be used for data generation in the data center and real-time use on a robot.

    Limits: Each size has its own hardware targets; larger models need data center GPUs.

  • Post-training with Cosmos Framework2

    Open framework for supervised fine-tuning, LoRA, distillation and reinforcement learning post-training; TAO 7 adds agent skills for fine-tuning.

    Why it matters: Adapts the base model to specific cameras, robot bodies and tasks.

    Limits: Needs curated domain data and GPU time measured in hours or more.

  • Data curation and evaluation2

    Cosmos Curator filters, annotates and deduplicates large sensor datasets; Cosmos Evaluator scores generated outputs at scale.

    Why it matters: Data quality and automatic scoring decide whether synthetic data helps.

    Limits: Automated scores are a filter, not a guarantee of usefulness for training.

  • Flexible serving backends1

    Runs on Diffusers, Transformers, vLLM, vLLM-Omni, SGLang, TensorRT LLM and NVIDIA NIM.

    Why it matters: Fits existing inference stacks, including OpenAI-compatible APIs.

    Limits: Backend support and performance vary by model and modality.

  • Default guardrails1

    Generation ships with guardrails turned on by default.

    Why it matters: Reduces unwanted outputs when generating content.

    Limits: Access to the guardrail model is gated on Hugging Face and must be requested.

Practical use cases

An autonomous vehicle dataset lacks rain, night and regional variety.
Approach
Use the Generator to vary existing drives with new weather, lighting and locations, then curate and evaluate the results before training.
Role of NVIDIA Cosmos
Generates and transforms the sensor video.
Data, infrastructure and skills
Existing drive recordings, data center GPUs and a validation set of real data.
Type of benefit
Wider data coverage
Caveats
Generated video can contain artifacts; never use it as the only test evidence.
First step
Run the autonomous driving inference recipe from the Cosmos Cookbook on one clip.
A robot needs a policy for a new task and body without training from zero.
Approach
Post-train Cosmos 3 on embodiment-specific camera and action data with Cosmos Framework, then test in closed-loop simulation.
Role of NVIDIA Cosmos
Acts as the pretrained backbone for the policy model.
Data, infrastructure and skills
Demonstration data from the target robot and GPU capacity for post-training.
Type of benefit
Faster policy development
Caveats
Policies need real-world testing and safety analysis.
First step
Review the Cosmos3-Edge policy example and its supported embodiments.
Operators want to ask questions about live or recorded camera video.
Approach
Use the Reasoner as the vision language model in a Metropolis VSS video analytics agent.
Role of NVIDIA Cosmos
Provides captions, alerts and answers about video content.
Data, infrastructure and skills
Camera streams, GPUs for inference and a defined alert policy.
Type of benefit
Faster review of video
Caveats
Check the license terms of the specific model build used inside each blueprint.
First step
Run the Reasoner notebook on a sample of your own footage.

Works with

Optional integration

  • NVIDIA NIMCosmos 3 can be served through NVIDIA NIM microservices.

Complementary tools

  • NVIDIA OmniverseNVIDIA describes Omniverse as the simulation environment and Cosmos as the models that turn simulations into photoreal synthetic data.
  • NVIDIA IsaacIsaac Sim and Cosmos together generate synthetic data for robot perception.
  • NVIDIA MetropolisCosmos Reason and Cosmos-Embed models run inside the Metropolis VSS blueprint.

Optional integration for

  • NVIDIA IsaacCosmos models can augment synthetic data generated in Isaac Sim.
  • NVIDIA MetropolisVSS uses Cosmos Reason and Cosmos-Embed models for captions, alerts and search.

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Try a hosted model

    Use the Try Now link on the Cosmos product page to test video generation and visual reasoning without installing anything.

    Check: You can judge whether output quality fits your use case.

  2. Request guardrail access

    Request access to nvidia/Cosmos-1.0-Guardrail on Hugging Face and log in with a read token from the same account.

    Check: Access is granted on the model page.

  3. Generate a first video locally

    Create a Python 3.13 environment with uv, install diffusers and dependencies, and run Cosmos3-Nano with the README quickstart prompt.

    Check: first_video.mp4 is written to disk.

  4. Reason over your own footage

    Open the Reasoner notebook in the cookbooks folder and run it on a short clip from your cameras.

    Check: Answers are correct for clips you have already labeled.

  5. Post-train and evaluate

    Follow a Cosmos Cookbook recipe with Cosmos Framework, then score results with the evaluation suites.

    Check: Scores on your validation set improve over the base model.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA Cosmos as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA Cosmos repository README (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  2. NVIDIA Cosmos (product page and FAQ) (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  3. NVIDIA VSS Blueprint documentation: License Information (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)