At the CVPR (Computer Vision and Pattern Recognition) conference in Denver, CO, today, and at GTC Taipei and Computex in Taiwan earlier this week, Nvidia announced new technology advances and global partnerships meant to give AV (autonomous vehicle) developers the AI building blocks required for safe Level 4 autonomy in the creation of next-generation, at-scale robotaxis, shuttles, passenger cars, and trucks.

At the IEEE Computer Society‘s CVPR, Nvidia is unveiling new physical AI agent skills that help researchers and developers speed the development of AVs, robots, and vision AI systems. At GTC and Computex, the major AV announcements included Drive Hyperion adoptions by Foxconn, Autobrains, Vinfast, Uber, and Humain, underscoring the company’s position as a common foundation for safe, scalable robotaxis and other AVs globally. The AI powerhouse also launched Alpamayo 2 Super, which it says is the world’s largest open reasoning model for robotaxis, which, combined with AlpaGym, OmniDreams, and new Omniverse NuRec libraries, helps developers build better Level 4 autonomy systems.

All of this comes as the leading-edge AI industry has evolved from VLMs (vision-language models), adding action headers for VLA (vision-language-action) models, and on to WAMs (world action models), according to Spencer Huang, Director of Product for Robotics at Nvidia.

“We realized that we needed to have spatial intelligence, so we went to world models,” he said. “Then we realized that you have to have both vision and action to be first-class citizens together, and so that becomes the world action model. That’s really where the world is heading.”

 

Enabling physical AI research with agent skills

At CVPR, Nvidia is unveiling new physical AI agent skills that help researchers and developers speed the development of autonomous vehicles, robots, and vision AI systems.

According to a preview of an Nvidia blog provided to Futurride about the agentic AI, company experts believe the core challenge in physical AI research is not only developing stronger models but also building a full workflow around them—reconstructing real-world scenes, generating edge-case scenarios, training policies, evaluating behavior, and rapidly iterating. These steps are currently fragmented across separate tools, slowing the pace of experimentation as researchers struggle to piece them together.

The company’s physical AI skills help AI agents work across Nvidia libraries, simulation frameworks, and open models such as Cosmos 3—for vision reasoning as well as world and action generation—to automate and scale an end-to-end loop at a faster pace.

For AV researchers, the problem is the “long tail” of driving or rare interactions, unusual road geometry, lighting changes and edge-case behaviors that are difficult to repeatedly collect, but critical for training and validation.

With Nvidia’s autonomous vehicle skills, researchers and developers can task AI agents to automate workflows for scene reconstruction from fleet data and generate synthetic scenarios. Neural reconstruction skills help AI agents turn fleet-captured data into editable 3D scenes for simulation and synthetic data generation, while technologies including Nvidia’s Omniverse NuRec, InstantNuRec, Harmonizer, and HiGS accelerated renderer help accelerate reconstruction, improve scene realism, and generate new views.

For AV researchers, repeatable simulation helps vary conditions, compare system responses, and uncover failure modes across scenarios beyond what can be captured in real-world data. Nvidia AlpaGym, an open-source closed-loop reinforcement learning framework, extends that approach by connecting policy rollouts and high-fidelity simulation with agent skills, scaling across thousands of GPUs, to help researchers move through setup, rollout, and evaluation.

Nvidia is also advancing AV research with its most powerful open driving foundation model to date: Nvidia Alpamayo 2 Super, an open 32-billion-parameter reasoning VLA model that reasons, plans, and acts across the full driving stack for safer, scalable level 4 development and deployment.

 

Research unlocking smarter autonomous driving

What makes an AV system safe is not only that it can reason through a situation but also that it can do so quickly enough on the hardware installed in the car, according to an Nvidia blog post.

At this year’s CVPR, Nvidia Research is presenting a paper that addresses the challenges and shares that training at scale creates systems that better generalize across diverse applications. The company’s LCDrive concept introduces a model that replaces expensive text-based reasoning with compact latent representations, letting autonomous vehicles think faster on embedded hardware.

In recent years, researchers have found that letting an AI reason—generating intermediate thinking steps before committing to an answer—reliably improves its decision-making. For AVs, the challenge is doing that reasoning on the hardware inside a vehicle. Text-based, chain-of-thought reasoning generates words, and every word is a token that takes time to produce. On the processor running inside a car, token count is a real constraint on how fast the system can respond.

LCDrive tackles this problem by replacing words with compressed latent representations.

Instead of generating human-readable reasoning steps, the system thinks in a compact latent space—states that capture spatial information rather than producing text. The architecture alternates between two kinds of thinking: proposing candidate actions, then predicting what the world will look like if those actions are taken.

It uses that predicted world state to refine its next step. It’s the same reasoning loop, just in a more computationally efficient form than natural language. The result is comparable output trajectory quality to text-based reasoning, using roughly half the tokens.

The model was built on Nvidia’s Alpamayo and trained using supervision derived from existing vehicle data.

 

Drive Hyperion gets more robotaxi traction

Nvidia announced a major expansion of its Drive Hyperion platform ecosystem with ride-hailing mobility providers to build and expand Level 4-ready robotaxi fleets.

“Autonomous mobility is entering its industrial scaling moment,” said Jensen Huang, Founder and CEO of Nvidia. “Vehicles are becoming robots, and robotaxi fleets will require AI infrastructure that can perceive, reason, and operate safely in the real world. Nvidia Drive Hyperion gives the world’s automakers, AV developers, and mobility networks a common Level 4-ready foundation—uniting compute, sensors, safety software and a global ecosystem to bring robotaxis from pilots to everyday transportation at scale.”

Built on the Nvidia Halos full-stack safety system for physical AI, Drive Hyperion combines high-performance Nvidia Drive AGX in-vehicle compute, the  Halos OS software foundation built on the safety-certified DriveOS operating system, with a compatible multimodal sensor suite and Drive AV software for highly automated and autonomous driving capabilities.

Among the collaborating automakers, Tier 1 suppliers, software and mobility providers in Asia, Europe and the Middle East to scale robotaxi deployments.

Foxconn is expanding its strategic collaboration with Nvidia to accelerate the development and planned deployment of Level 4-ready robotaxi fleets. The effort combines Foxconn’s contract design and manufacturing services with Nvidia’s Drive Hyperion platform to support the rapid integration, scaling, and deployment of Level 4 EVs (electric vehicles), starting in Taiwan, with Kaohsiung expected to serve as an early deployment city, then expanding across Asia.

“Autonomous mobility is a strategic focus of Foxconn’s EV initiative,” said Young Liu, Chairman of Foxconn. “By leveraging strategic partnerships and Nvidia’s capabilities, we are accelerating the deployment of level 4 robotaxi technology, with Foxconn providing high-performance computing and sensor integration to enable a worldwide rollout across communities and cities.”

The collaboration reinforces Taiwan’s role in the global autonomous mobility ecosystem, building a scalable, safety-focused robotaxi platform to advance more efficient urban transportation. Foxconn plans to launch a robotaxi service in 2028, starting with airport-to-city routes and later expanding along corridors linked to Taiwan’s high-speed rail network.

“The collaboration between Foxconn, Foxtron, and Nvidia represents an important milestone in accelerating Taiwan’s transformation into a world-class smart city ecosystem,” said Chen Chi-mai, Kaohsiung’s Mayor. “As Taiwan’s major industrial and innovation hub, Kaohsiung is actively investing in smart infrastructure, green mobility and AI-driven urban development.”

VinFast is working with Autobrains to bring Level 4 vehicles built on Drive Hyperion to Southeast Asia, combining the automaker’s vehicle development and manufacturing capabilities with Autobrains’ autonomous driving software stack.

“Advanced mobility shouldn’t be a luxury,” said Duong Nguyen, Deputy CEO of ADAS at VinFast Global. “VinFast is committed to building scalable and accessible autonomous driving solutions through collaboration with global technology leaders. Together with Autobrains and Nvidia, we are exploring a practical and cost-efficient path toward Level 4 mobility for Southeast Asia’s highly dynamic real-world traffic environments.”

In Europe, Uber is also working with Autobrains to launch a robotaxi program in Munich built on Drive Hyperion, integrating Autobrains’ agentic AI autonomous driving software to support scalable, Level 4-ready robotaxi operations.

“Autonomous driving will not scale by relying on a single model to solve every driving scenario,” said Igal Raichelgauz, Founder and CEO of Autobrains. “It requires systems that can reason, adapt, and make decisions under uncertainty. With Uber and Nvidia, we are bringing Autobrains’ Agentic AI into autonomous ride-hailing—combining reasoning-based driving intelligence with the mobility platform and automotive compute needed to support scalable robotaxi operations across cities, vehicles, and real-world conditions.”

The effort expands Uber’s growing reach in the European ride-hailing market, with the automaker involved to be announced later this year.

“For automakers and autonomy developers, the challenge is not just building autonomous vehicles; it’s bringing them into a commercial network where they can reliably serve riders at scale,” said Sarfraz Maredia, Global Head of Autonomous Mobility and Delivery at Uber. “This program creates a new path to do that by combining vehicle-agnostic autonomy, leading AI compute, and Uber’s ride-hailing platform.”

Humain is working to bring Drive Hyperion-powered robotaxis to the Middle East, expanding the platform’s regional footprint and leveraging its AI and mobility ecosystem to support the development and deployment of Level 4-ready autonomous transportation solutions across the region.

“Autonomous mobility will become one of the defining AI platforms of the next decade,” said Tareq Amin, CEO of Humain. “By working with Nvidia, Humain is helping enable the infrastructure, intelligence, and operational scale needed to develop and support the future of Level 4-ready transportation in Saudi Arabia. This collaboration reflects our broader vision to help build AI-native infrastructure platforms that connect the digital and physical worlds at scale.”

 

Alpamayo 2 Super open reasoning model for robotaxis

At GTC and Computex, Nvidia introduced Alpamayo 2 Super, a 32-billion-parameter reasoning‑based VLA (vision-language-action) model that extends the family of open AI models, simulation frameworks, and physical AI datasets for safe, Level 4 robotaxi development.

“Alpamayo is the moment cars begin to safely reason, not just drive,” said Huang. “Only Nvidia makes available open models, simulation, real-world data, and agent skills so the entire global robotaxi ecosystem can develop Level 4 capabilities that understand edge cases, explain decisions, earn trust, and scale safely to millions of vehicles.”

Alpamayo 2 Super helps accelerate AV development by eliminating the need to build key autonomy infrastructure from scratch. It enables humanlike perception, reasoning, and action—and provides the interpretability needed for safety validation and regulatory collaboration.

Alongside the model, the company announced new tools, models, and agent skills that complete the pipeline from real-world data capture to closed-loop training and in-vehicle deployment, including AlpaGym, OmniDreams, and new Omniverse NuRec models.

To better train models for on-road deployment, the AlpaGym framework provides a platform for closed-loop reinforcement learning. The OmniDreams generative world model for photorealistic closed-loop AV scenario generation enables developers to simulate rare and long-tail driving scenarios at scale.

To amplify developer productivity, Nvidia is providing physical AI agent skills for all of its AV development tools. For example, the neural reconstruction skill powered by Omniverse NuRec uses real-world fleet driving scenarios for simulation and generates synthetic training data at scale.

With Alpamayo 2, the Nvidia family is intended to go beyond trajectory generation to reason, plan, and act across the full driving stack. With multitask capabilities spanning reasoning, auto-labeling, scene understanding, model critiquing, and distilling knowledge into smaller models, it provides the building blocks for scalable Level 4 AV development and deployment.

Built on Nvidia Cosmos world foundation models, Alpamayo 2 Super scales from 10 billion to 32 billion parameters compared to improv reasoning, 3D spatial understanding, and trajectory prediction in long‑tail scenarios.

Its scope expands from front-focused cameras to 360-degree situational awareness across front, side, and rear views, giving the model a more complete context for safer lane changes, merges, and intersection crossings.

The model adds meta-action output, such as yield, lane change, and stop, so the model predicts high-level driving decisions for downstream planning in addition to trajectories and chain-of-causation (CoC) traces.

It introduces reasoning auto‑labeling with 2D grounding so the 32-billion-parameter foundation model can provide high-quality reasoning labels, compressing annotation cycles from months to days and reshaping AV data pipeline economics.

The model features improved CoC (chain of causation) traces and trajectories, especially in rare, complex, long‑tail scenarios where traditional imitation‑learning AV stacks struggle.

These advancements make Alpamayo 2 Super Nvidia’s most powerful open driving foundation model to date. Designed as a teacher model, it can be distilled into compact models that run on the accelerated compute of the Drive Hyperion platform and Drive AGX Thor running in the vehicle. With its 32 billion parameters, a downstream AV stack built on Alpamayo inherits higher‑quality reasoning and perception from a single open release without requiring each manufacturer to rebuild from scratch.

Alpamayo 2 Super is expected to be available this summer on GitHub for inference code and Hugging Face for model weights.

Since its launch, Alpamayo has been downloaded close to 400,000 times. The open platform includes post-training scripts that allow researchers and developers to adapt the models to their own datasets, scenarios, and driving policies.

 

AlpaGym for closed-loop training and deployment cycles

While open‑loop training evaluates models against recorded data and generates a single round of actions, AlpaGym runs models through continuous decision and observation cycles in Nvidia’s AlpaSim, with every braking, steering, and navigation choice affecting the environment. As a result, it exposes the compounding errors and edge‑case failures that static datasets miss, allowing models to learn from experience.

Built on the AlpaSim microservice simulation stack and Omniverse NuRec, AlpaGym enables efficient, scalable, closed-loop RL (reinforcement learning) to push the frontier of driving performance. In combination with the physical AI AV dataset, Alpamayo provides a continuous path from open-loop pretraining to closed-loop refinement.

Nvidia is also releasing the CoC auto-labeling pipeline as an open-source solution on GitHub. The pipeline automatically generates decision-grounded and causally linked CoC labels from raw driving clips with no human annotation required, providing the causal training data foundation needed to train embodied reasoning models at scale.

To support reasoning-based AV development, Nvidia is launching new physical AI agent skills as part of an agent toolkit to guide developers and their coding agents through the simulation, data generation, and closed-loop training workflows needed to build and validate autonomous driving systems at scale. This includes neural reconstruction skills powered by Omniverse NuRec libraries, OmniDreams skills for photorealistic scenario generation, and AlpaGym skills for closed-loop RL.