AWS Physical AI Blog
NVIDIA Cosmos 3 on AWS: Omnimodal World Models for Physical AI
Introduction
In our previous blog, we demonstrated how NVIDIA Cosmos world foundation models on AWS address the data scarcity problem in physical AI: the challenge of training robust perception, prediction, and control models when real-world robot and sensor data is expensive, unsafe, or impractical to collect at scale. We introduced two production-ready architectures: Cosmos NIM on Amazon EKS for real-time inference, and Cosmos containers on AWS Batch for high-throughput offline synthetic data generation.
Cosmos 3, released in May 2026 unifies the separate prediction, transfer, and reasoning models from earlier Cosmos releases with a single unified model. NVIDIA describes it as the first fully open omni-model for physical AI that reasons the physical world, generates photorealistic simulations, and produces executable robot actions in a single forward pass.
In this post, you’ll learn:
- How the Cosmos 3 Mixture-of-Transformers (MoT) architecture works internally
- Reference architectures for curation, post-training, inference, and edge deployment on AWS
- Instance sizing and parallelism strategies
- Where to find deployment resources and sample code to get started
What is NVIDIA Cosmos 3
Cosmos 3 is a frontier omni-model that unifies vision reasoning, world generation, and action prediction into a single architecture. Earlier Cosmos releases required orchestrating Cosmos Predict, Cosmos Transfer, and Cosmos Reason as three separate models, each with an independent inference call, its own model-loading overhead, and its own serving infrastructure. Cosmos 3 delivers all three capabilities through one model and one inference call.
Core Capabilities
| Capability | Description | Physical AI Application |
|---|---|---|
| Vision Language Model (VLM) | Understands and reasons across text, image, video, audio | Scene understanding, anomaly detection, quality inspection |
| World Model | Generates physically accurate future world states as video | Synthetic training data, scenario simulation, what-if analysis |
| World Action Model | Produces executable action trajectories (joint angles, gripper positions) | Robot policy training, manipulation planning |
| Forward Dynamics | Predicts future observations conditioned on robot actions | Model-based RL, trajectory optimization |
| Inverse Dynamics | Infers actions from observed demonstrations | Learning from video demonstrations, imitation learning |
| Policy Model | Predicts action sequences from observations + task prompts | End-to-end robot control |
Model Variants
| Variant | Parameters | Typical Use Case |
|---|---|---|
| Cosmos 3 Super | 64B (32B Reasoner + 32B Generator) | Highest-fidelity simulation, research, offline generation |
| Cosmos 3 Nano | 16B (8B Reasoner + 8B Generator) | Production inference, balanced cost/quality |
| Cosmos 3 Edge | 4B | On-device robot inference, low-latency control loops |
A New Architecture for Physical AI
Traditional physical AI pipelines chain multiple specialized models in sequence: a vision-language model reasons about a scene and produces a text description; that description conditions a separate video generation model to synthesize future frames; those frames are then passed to an independent policy network that outputs motor commands. Each handoff between models loses contextual information, adds inference latency, and increases engineering complexity in production systems.
The Cosmos 3 Mixture-of-Transformers (MoT) architecture replaces this multi-model pipeline with a single forward pass through shared parameters. A Reasoner tower encodes the scene as causally attended autoregressive (AR) tokens. A Generator tower simultaneously denoises video and action latents using bidirectional attention, conditioned on the Reasoner tower’s internal representations. One inference call produces a physically coherent video prediction along with the corresponding end-effector trajectory and gripper actuation commands, without an intermediate text bottleneck or cross-model handoff.
Figure 1: NVIDIA Cosmos 3 unified Mixture-of-Transformers (MoT) architecture.
Figure 1: Cosmos 3 Unified MoT Architecture
Example: Warehouse Robot Picking: Consider a warehouse robot picking irregularly shaped packages. In a traditional pipeline, the VLM might describe the scene as “a deformed cardboard box on a conveyor belt”, but spatial details such as the exact surface curvature and distance to neighboring objects are lost in that text description. The downstream video model then generates a plausible but imprecise future state, and the policy network operates on this degraded information. Cosmos 3 instead perceives the box geometry, reasons about grasp stability, generates the pick trajectory, and outputs motor commands in a single forward pass, preserving full spatial resolution throughout.
Dual-Tower Design
Cosmos 3 introduces a fundamentally new architecture that merges autoregressive reasoning and diffusion-based generation into a single, jointly trained model.
Figure 2: Cosmos 3 Deep Dive into MoT Architecture
1. Autoregressive Tower (Reasoning)
Multimodal inputs split into two paths. Text and ViT-encoded image patches form the AR subsequence routed to the Autoregressive tower. Video, audio, and action data pass through frozen VAE encoders into compact latents, forming the sequence routed to the Diffusion tower. 3D Multimodal RoPE (MRoPE) aligns all tokens onto a shared physical temporal axis with FPS modulation to handle different frame rates.
2. Autoregressive Tower (Reasoning)
The ViT encoder converts images into patch embeddings with DeepStack multi-layer aggregation. These visual tokens concatenate with language tokens into a single sequence. Causal self-attention processes left-to-right where each token attends only to prior tokens, building an incremental understanding of spatial relationships, object states, and physics. The next-token prediction layer outputs dense representations encoding the model’s complete scene reasoning.
3. Diffusion Tower (Generation)
The frozen VAE compresses video, audio, and actions into a unified latent space. Bidirectional attention processes all tokens globally, ensuring spatial and temporal coherence across the generated output. Iterative denoising progressively refines latent from noise into high-fidelity video, synchronized audio, and precise action trajectories.
4. Dual-Stream Joint Attention
Diffusion tokens compute attention over concatenated Keys/Values from both the Diffusion tower and the Autoregressive tower. Every generated pixel, audio sample, and action coordinate is directly conditioned on the model’s semantic reasoning. AR tokens remain causally self-contained, preserving reasoning integrity while sharing internal state with the generation process.
5. Unified Output
A single forward pass produces text reasoning, photorealistic video, synchronized audio, and action trajectories (joint angles, gripper commands, 6-DoF paths) simultaneously. No model chaining, no intermediate API calls, no orchestration infrastructure.
Reference Architecture: Cosmos 3 on AWS
The following architecture supports the full lifecycle, from data curation to model deployment to edge inference, using Cosmos 3 on AWS.
- Step 1 Data Curation: Raw video, LiDAR, and camera feeds land on Amazon S3. Cosmos Curator runs on AWS Batch to filter, de-duplicate, and quality-score the dataset. Cosmos Dataset Search deploys on Amazon EKS with Milvus vector DB and Cosmos-embed NIM for semantic retrieval of relevant training clips.
- Step 2 Post-Training: Domain adaptation (SFT) runs on Amazon SageMaker HyperPod to fine-tune Cosmos 3 for custom environments and tasks. Action post-training adapts the action head to specific robot embodiments, including joint configurations, gripper mappings, and 6-DoF end-effector paths. Cosmos 3’s built-in forward dynamics validates learned policies through closed-loop testing before physical deployment.
- Step 3 Production Inference: Two deployment options serve different workload patterns: Amazon EKS with Cosmos NIM for real-time inference and managed inference with Amazon SageMaker JumpStart.
- Step 4 Edge Deployment: Cosmos 3 Edge (4B) deploys NVIDIA Jetson hardware via AWS IoT Greengrass for real-time on-device inference. OTA model updates push improved models through a continuous deployment pipeline. Edge sensor data flows back to S3, closing the loop for retraining and continuous improvement.
Security and Governance
- The architecture enforces SSE-KMS encryption on all S3 buckets at rest. We configure VPC endpoints for S3/ECR to eliminate public internet egress.
- HyperPod clusters and EKS node groups run in private subnets, and IAM roles are scoped per pipeline stage (curation, training, inference, edge) rather than a single broad role.
- Edge devices authenticate to IoT Greengrass via device certificates. We rotate certificates on a defined schedule and restrict OTA update channels to signed model artifacts only.
- For customer or safety-critical robot data, the architecture uses AWS PrivateLink for inter-service traffic and enables CloudTrail logging on all model artifact promotions.
EC2 Instance Recommendation
Instance sizing depends on the Cosmos 3 runtime surface. The full Cosmos3-Nano and Cosmos3-Super checkpoints are currently tested in BF16. The Cosmos3 Generator NIM supports separate FP8 profiles with different memory and parallelism requirements.
| Workload | Recommended instance | Guidance |
|---|---|---|
| Nano single-GPU inference | g7e.2xlarge or p5.4xlarge | One RTX PRO 6000 96 GB or H100 80 GB. Use BF16 for the full checkpoint; FP8 may be used with supported NIM profiles. |
| Nano throughput or post-training | p5.48xlarge | Eight H100 GPUs. Supports scale-out inference and matches NVIDIA’s tested eight-GPU SFT configuration. |
| Super full-model inference | p5e.48xlarge or p5.48xlarge | Use the runtime-recommended HSDP, CFG, and Ulysses layout rather than assuming TP8. H200 provides more memory headroom. |
| Multi-node post-training | p5en.48xlarge | Eight H200 GPUs with EFAv3. Suitable for NCCL collectives, but benchmark against the NVIDIA-tested H100 recipe. |
| Blackwell inference or training | p6-b200.48xlarge | Eight B200 GPUs with 1,432 GiB aggregate GPU memory. Validate the runtime’s sharding or replication layout before making performance claims. |
Deployment of Cosmos 3 on AWS
Cosmos 3 supports two production deployment patterns on AWS. Choose based on your operating model: Amazon EKS with Cosmos NIM gives platform teams full control over the serving stack for latency-sensitive, high-throughput workloads, while Amazon SageMaker JumpStart offers a fully managed path from model catalog to callable endpoint with no infrastructure to operate.
Pattern 1: Real-Time Inference on Amazon EKS with Cosmos NIM
For low-latency workloads such as robotics control loops, real-time scene understanding, and interactive simulation. Cosmos NIM microservices are pulled from NVIDIA NGC (or mirrored through Amazon ECR) and deployed as GPU-backed pods on Amazon EKS. Karpenter provisions right-sized GPU nodes on demand and scales the fleet with load. Multi-AZ node placement provides high availability. A new NIM image version deploys alongside the old, passes health checks, and takes traffic before the prior revision drains. Pods assume scoped IAM roles via IRSA for S3 access to prompt assets and generated outputs, and inference metrics (GPU utilization, queue depth, per-request latency) flow to Amazon CloudWatch and Prometheus to drive horizontal pod autoscaling. This pattern suits teams already operating Kubernetes platforms who need control over batching, routing, and GPU bin-packing.
Figure 3: Real-time inference with Cosmos 3 on Amazon EKS
Pattern 2: Managed Inference with Amazon SageMaker JumpStart
For teams that want production inference without managing infrastructure. As of August 2026, all three Cosmos 3 variants – Edge, Nano, and Super – are available in SageMaker JumpStart and deploy in a few clicks from the model catalog in the SageMaker console, or programmatically through the SageMaker Python SDK. Real-time endpoints serve interactive requests with built-in auto scaling against traffic, while asynchronous endpoints handle long-running world-generation jobs without client-side timeouts. The pattern also closes the loop with post-training: checkpoints produced on SageMaker HyperPod and registered in the SageMaker Model Registry deploy to the same managed endpoints, keeping fine-tuning and serving inside a single managed service. Endpoints run inside your VPC with IAM-scoped invocation; SageMaker handles provisioning, health checks, and rolling model updates behind a single HTTPS API, and invocation metrics stream to CloudWatch. This is the lowest-operational-overhead path from model selection to production inference.
Figure 4: Managed Cosmos 3 inference with Amazon SageMaker JumpStart
Why AWS for Cosmos 3 Workloads
- Broadest NVIDIA GPU portfolio: AWS offers P5 (H100), P5e (H200), P6 (B200/B300), P6e (GB200 UltraServers with 72 GPUs in a single NVIDIA NVLink domain), and G7e (RTX PRO 6000) instances.
- Production-validated deployment patterns: AWS publishes official reference architectures for Cosmos NIM on Amazon EKS and Cosmos containers on AWS Batch. SageMaker JumpStart provides one-click deployment of all three Cosmos 3 variants, and SageMaker HyperPod supports multi-node training with auto-recovery node failures.
- End-to-end pipeline support: The full Cosmos lifecycle runs natively on AWS without external dependencies – from S3 ingestion through Batch curation, EKS-based semantic retrieval, HyperPod post-training, and IoT Greengrass edge deployment.
- Cost optimization at scale: AWS reduced GPU instance pricing by up to 45% in June 2025. Spot Instances deliver up to 90% savings for batch synthetic data generation workloads. Capacity Blocks reserve GPU capacity for predictable fine-tuning windows without on-demand pricing uncertainty. Karpenter scales GPU node capacity to zero during idle periods, eliminating GPU cost when demand drops.
Conclusion
NVIDIA Cosmos 3 simplifies the Physical AI stack by combining scene understanding, world generation, and action prediction in a single model. Instead of stitching together multiple specialized models and inference paths, teams can use one architecture to reason about the environment, simulate future states, and generate executable robot actions.
On AWS, that model maps to an end-to-end deployment workflow: curate and store data in Amazon S3, adapt the model with SageMaker HyperPod, serve it with Amazon EKS and Cosmos NIM or through SageMaker JumpStart, and extend it to edge devices with AWS IoT Greengrass. Together, these patterns help teams move from experimentation to production with less orchestration overhead and a clearer path to scaling Physical AI workloads.
To get started, use the reference architectures and sample code linked below and choose the deployment pattern that best fits your workload: EKS for maximum control, JumpStart for the fastest managed path.