NVIDIA

NVIDIA AI News & Updates60 Updates

The latest AI news and updates of NVIDIA — AI computing company building GPUs, inference hardware, and developer platforms for training and deployment. Covering NVIDIA's latest company news, research, and product updates from the past 90 days.

NVIDIANVIDIASep 10

NVIDIA and USC Unveil HorizonRelight for Consistent Long-Horizon Video Relighting

NVIDIA Research and USC released HorizonRelight, a method for relighting long videos consistently. By propagating target-domain context across sliding-window chunks, the system prevents visible lighting jumps and boundary artifacts common in chunked inference. The research, presented at ECCV 2026, includes a demo and paper, with code release coming soon.

Read more
NVIDIANVIDIASep 10

NVIDIA Details Encode-Prefill-Decode Disaggregation for Multimodal Model Serving

NVIDIA published guidance on using encode-prefill-decode (EPD) disaggregation in its Dynamo inference framework to separate vision encoding from LLM prefill and decode stages. This technique reduces resource contention and improves response times for image-heavy workloads, delivering up to 5x faster time to first token and 7x faster end-to-end latency in specific configurations.

Read more
NVIDIANVIDIASep 10

NVIDIA and Palantir Combine AI Stacks for Supply Chain Decision Intelligence

NVIDIA and Palantir are integrating NVIDIA Nemotron, cuOpt, and NeMo with Palantir Foundry and AIP to optimize supply chain operations. This AI stack, first deployed within NVIDIA’s own network, uses post-trained open models to flag emerging risks and model supply constraints. The system generates actionable recommendations and trade-off analysis while maintaining control over proprietary data.

Read more
NVIDIANVIDIASep 10

NVIDIA Partners With Australian Ecosystem to Build 2-Gigawatt AI Factories

NVIDIA is collaborating with Australian data center providers, including Firmus, CDC, and AirTrunk, to expand land, power, and shell capacity for AI factories. Built on the NVIDIA DSX platform, these facilities provide high-performance compute for local researchers and enterprises. The partnership aims to deliver up to a 2-gigawatt buildout by 2027 to meet surging regional demand.

Read more
NVIDIANVIDIASep 5

NVIDIA Nemotron Model Outscores Top Human Coders at IOI 2026

NVIDIA’s fine-tuned Nemotron model scored 535.4 out of 600 on the IOI 2026 problem set, exceeding both the gold medal threshold and the top human contestant’s score. The model competed unofficially in Uzbekistan under identical constraints as human participants, including no internet access and strict time limits, marking the first AI system to outscore the highest-scoring human contestant.

Read more
NVIDIANVIDIASep 4

Jim Fan Reflects on the Decade-Long Path to GPT-6 Astra

Jim Fan, co-creator of OpenAI's 2016 World of Bits project, notes that GPT-6 Astra finally achieves reliable computer use, a task that previously doomed RL-from-scratch approaches. He attributes this success to the Specialized Generalist recipe—broad pretraining followed by specialization to screen pixels—and the massive compute required to produce the emergent capability.

Read more
NVIDIANVIDIASep 3

NVIDIA to Acquire Hugging Face for $12.93 Billion

NVIDIA entered a definitive agreement to acquire Hugging Face for $12.93 billion. The platform will remain an open, neutral, and independent home for the AI ecosystem, with the founders and team continuing their mission. NVIDIA plans to scale the platform’s infrastructure while maintaining support for multi-cloud, multi-accelerator development and open-weight models from all builders.

Read more
NVIDIANVIDIASep 1

NVIDIA and CrowdStrike Evaluate Agentic System for Adaptive Cyber Defense

NVIDIA and CrowdStrike evaluated an agentic cybersecurity system that connects offensive and defensive agents in a continuous loop. Using Nemotron 3 Ultra for orchestration and a fine-tuned Nemotron 3 Super for detection generation, the pipeline improved backtest detection rates to 41.9%. In live-fire testing, the system generated three detection rules that successfully caught eight previously unseen attacks.

Read more
NVIDIANVIDIAAug 31

NVIDIA and MediaTek Deepen Partnership Across AI Infrastructure and Automotive

NVIDIA and MediaTek are expanding their collaboration to build next-generation AI computing platforms, supported by a 3.5 billion dollar NVIDIA investment in MediaTek. MediaTek will adopt the NVIDIA NVLink Fusion platform to integrate custom XPUs into rack-scale AI factories. The companies are also continuing joint development of RTX Spark PC chips and AI-powered, software-defined vehicle platforms.

Read more
NVIDIANVIDIAAug 27

NVIDIA and AWS Expand Full-Stack Partnership for Agentic and Physical AI

NVIDIA and AWS are expanding their partnership to deploy 2 million additional GPUs across AWS global infrastructure in 2027-2028. The collaboration adds NVIDIA Vera CPU-based infrastructure, extends NVLink Fusion with custom high-bandwidth memory, and builds AI factories for the U.S. government, including 100,000 GPUs for secure federal workloads. NVIDIA reports that demand is currently exceeding all previous forecasts.

Read more
NVIDIANVIDIAAug 26

NVIDIA Dynamo Adds Shadow Engine Recovery for Near-Instant LLM Failover

NVIDIA introduced shadow engine recovery as a preview feature in its Dynamo inference framework. By maintaining a pre-warmed standby engine on the same GPU, the system eliminates the need for cold restarts during process failures. Tests with GLM-5.2 on B200 nodes restored serving capacity in 7.3 seconds, nearly 39 times faster than a standard cold restart.

Read more
NVIDIANVIDIAAug 25

NVIDIA Ships Groq 3 LPX Accelerator for Ultrafast Agentic AI

NVIDIA has moved its Groq 3 LPX interactive AI inference accelerator into full production. Designed for the Vera Rubin NVL72 platform, the accelerator delivers 3,400 output tokens per second on Gemma 4 31B, providing 4x faster responsiveness for agentic workloads. Nebius is the first AI cloud to adopt the platform, with Groq also planning early adoption.

Read more
NVIDIANVIDIAAug 25

NVIDIA Nemotron 3.5 Lightning Ranks Top 4 on Agentic Benchmarks

NVIDIA Nemotron 3.5 Lightning ranked in the top 4 open-weight models on the PinchBench leaderboard, achieving an 86.4% average success rate on OpenClaw agent tests. The 30B parameter mixture-of-experts model is optimized for high-volume execution. NVIDIA’s Nemotron 3 Ultra maintains the #1 spot on the same benchmark, highlighting the family's performance across agentic tasks.

Read more
NVIDIANVIDIAAug 25

xAI Deploys NVIDIA Vera CPUs to Accelerate Agentic AI Workloads

NVIDIA announced that xAI is deploying Vera CPUs to accelerate its agentic AI workloads. The Vera CPU, featuring 88 Olympus cores and 1.2TB/s bandwidth, delivers up to 1.8x faster task completion than x86 alternatives. xAI is also expanding its Grok infrastructure on the Vera Rubin platform and plans to integrate optimized Vera Rubin NVL72 systems into its Starmind satellites.

Read more
NVIDIANVIDIAAug 22

NVIDIA and UC Berkeley Release T-Rex Tactile-Reactive Robotics Methodology

NVIDIA and UC Berkeley open-sourced T-Rex, a tactile-reactive manipulation framework that gives robots a dedicated high-frequency control loop for touch. The release includes a 50-hour tactile-synchronized robot play dataset on HuggingFace and a three-stage training recipe. The architecture uses a Mixture-of-Transformer model to process tactile corrections at four times the frequency of visual planning.

NVIDIANVIDIAAug 22

NVIDIA AVO Agent Scores 100% on ARC-AGI-3 Interactive Reasoning Benchmark

NVIDIA’s AVO agent architecture achieved a 100.00 RHAE score on the ARC-AGI-3 benchmark, completing all 183 levels across 25 environments. The system, which uses persistent memory and supervision to sustain long-horizon tasks, solved the benchmark using 6,624 environment actions. This result demonstrates that system-level architecture can effectively convert model capability into sustained autonomous progress.

Read more
NVIDIANVIDIAAug 22

NVIDIA cuOpt Ranks Fastest on Hans Mittelmann Optimization Benchmarks

NVIDIA cuOpt is now the fastest open-source solver on Hans Mittelmann benchmarks across three optimization problem classes: linear programming, second-order cone programming, and quadratic programming. The GPU-accelerated engine provides near real-time solutions for large-scale problems, integrating into existing modeling languages for deployment across hybrid and multi-cloud environments.

Read more
NVIDIANVIDIAAug 22

NVIDIA Benchmarks 300+ Verified Agent Skills Using Open-Source SkillEvaluator

NVIDIA benchmarked over 300 verified agent skills, finding they improved correctness by 41 points, effectiveness by 39, and efficiency by 35. The company open-sourced SkillEvaluator, a tool that measures skill impact through static checks and live task runs in isolated environments. This allows developers to test and verify skill performance before deployment.

Read more
NVIDIANVIDIAAug 22

NVIDIA Releases TensorRT Model Connect for Two-Command Model Deployment

NVIDIA released TensorRT Model Connect in public preview, an open-source tool that converts supported Hugging Face models into end-to-end TensorRT inference bundles in two commands. The tool eliminates intermediate ONNX export steps and provides bundles callable through native C++ APIs. NVIDIA built the project using OpenAI Codex agents, with human developers directing and reviewing the implementation.

NVIDIANVIDIAAug 22

NVIDIA Guarantees Exclusive AI Compute Capacity at Ohio Technology Campus

NVIDIA guaranteed SB Energy’s PORTS-Pike Technology Campus in Ohio to exclusively host NVIDIA AI compute. The site will use NVIDIA’s DSX platform to provide 8 IT-GW of capacity for OpenAI. NVIDIA is investing $1.5 billion in SB Energy to support the project, which begins phased operations in 2028 and aims to create tens of thousands of local jobs.

Read more
NVIDIANVIDIAAug 15

NVIDIA Launches NeMo Switchyard for Dynamic AI Agent Model Routing

NVIDIA launched NeMo Switchyard, an open-source library that dynamically routes AI agent tasks between models based on capability, cost, and latency. The system directs complex reasoning to frontier models while offloading high-volume execution to specialized models like Nemotron 3.5 Lightning. The routing layer is provider-agnostic and open-source — agent workloads stay optimized without rebuilding applications when switching or adding model providers.

Read more
NVIDIANVIDIAAug 14

NVIDIA and LG Expand Collaboration on Physical AI and Robotics

NVIDIA and LG are expanding their collaboration to accelerate physical AI development. LG plans to unveil a bipedal humanoid robot built on NVIDIA Isaac GR00T and Jetson Thor in the first quarter of next year. The partnership also includes validating wheel-based robots in Tennessee, building AI factory reference sites with NVIDIA Vera Rubin, and developing AI-defined vehicle platforms.

Read more
NVIDIANVIDIAAug 10

NVIDIA Partners With Six Financial Firms to Mobilize $500B for AI

NVIDIA signed memorandums of understanding with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to establish independent compute financing platforms. These partnerships aim to mobilize over $500 billion of third-party capital to fund the global buildout of AI infrastructure. The platforms create dedicated capital pools to help customers access large-scale compute for AI factories.

NVIDIANVIDIAAug 10

NVIDIA Optimizes Meta's Muse Glimmer for Local Agentic AI Workflows

NVIDIA announces support for Meta’s Muse Glimmer, a 30B open-weight dense model designed for long-running agentic tasks. The model runs locally on NVIDIA edge, desktop, and workstation platforms, delivering up to 20K tokens per second on a single GPU. Developers can deploy the model using NVIDIA NIM containers, SGLang, or vLLM inference recipes.

Read more
NVIDIANVIDIAAug 8

Firebird Launches Armenia AI Factory Powered by NVIDIA Accelerated Computing

Firebird launched the CIS region’s largest AI factory in Armenia, integrating NVIDIA accelerated computing with Dell PowerEdge servers and Schneider Electric power infrastructure. The facility utilizes the NVIDIA DSX platform to support AI training and deployment. Firebird plans to deploy over 70,000 NVIDIA Rubin and Blackwell GPUs and 300 megawatts of capacity by the end of 2027.

Read more
NVIDIANVIDIAAug 4

NVIDIA Joins NSF’s $100M State and Regional AI Infrastructure Hubs

NVIDIA is participating in the U.S. National Science Foundation’s new $100 million State and Regional AI Infrastructure Hubs program. The initiative expands access to advanced computing, data, software, and technical expertise for colleges and universities. These regional hubs aim to accelerate scientific research, strengthen AI education, and prepare students for the AI economy through shared resources and training.

Read more
NVIDIANVIDIAAug 3

NVIDIA Details Four Architecture Choices for Long-Context Inference Performance

NVIDIA published a technical guide on co-designing AI model attention to optimize inference speed for long-context workloads. The post identifies four architecture levers—group size, head dimension, KV-cache size, and parallelism strategy—that determine performance ceilings before training begins. It provides specific guidelines for aligning these choices with GPU hardware to maximize throughput and interactivity.

Read more
NVIDIANVIDIAJul 31

NVIDIA Research Releases Spatial-IQ Benchmark for 3D Spatial Reasoning

NVIDIA Research released Spatial-IQ, a diagnostic benchmark that decomposes 3D object counting into nine hierarchical sub-tasks. While humans reach 82.1% accuracy, top multimodal models manage 17.7% by bypassing the underlying spatial logic. Training a Qwen2.5-VL-32B model on these sub-tasks using chain-of-thought supervision and reinforcement learning improved its counting accuracy from 2.9% to 62.6%.

Read more
NVIDIANVIDIAJul 29

NVIDIA Launches Cosmos-Dreams Neural Simulator for Real-Time Physical AI

NVIDIA announced Cosmos-Dreams, a neural closed-loop simulator for physical AI, at SIGGRAPH 2026. The model generates interactive, physics-aware virtual worlds from a single frame, enabling autonomous systems to train in synthetic environments. The simulator runs in real-time on a single NVIDIA RTX Pro 6000 workstation GPU, shifting data collection to compute-driven scenario generation.

Watch
NVIDIANVIDIAJul 28

NVIDIA Cosmos Models Surpass 10 Million Downloads on Hugging Face

NVIDIA’s Cosmos world foundation models have reached 10 million downloads on Hugging Face. This milestone reflects widespread adoption among researchers and engineers building physical AI systems, including robots and autonomous vehicles. The open model platform provides the foundation for real-time scene understanding, reasoning, and action generation in memory-constrained edge environments.

Read more
NVIDIANVIDIAJul 27

NVIDIA Launches Open Secure AI Alliance and NOOA Agent Framework

NVIDIA introduced the Open Secure AI Alliance to develop industry-wide AI security standards. As a founding contribution, NVIDIA released the Nvidia Labs Object-Oriented Agents (NOOA) framework. This open-source research project structures agents as Python classes, achieving state-of-the-art performance on software engineering, cybersecurity, and reasoning benchmarks while reducing token costs through a typed, object-oriented harness architecture.

Read more
NVIDIANVIDIAJul 27

NVIDIA Nemotron 3 Ultra Leads Open Models in Agentic RTL Coding

NVIDIA tested its Nemotron 3 Ultra model on agentic RTL chip-design tasks using the ACE-RTL agent. The model achieved a 97.1% average pass rate across nine design categories on the CVDP benchmark. It maintained this accuracy while using 6,629 tokens per iteration, outperforming other open models in both pass rate and token efficiency.

Read more
NVIDIANVIDIAJul 24

NVIDIA ModelExpress Cuts DeepSeek-V4 Pro Startup Time to Under Two Minutes

NVIDIA launched ModelExpress to accelerate model weight distribution, reducing DeepSeek-V4 Pro startup time from 8 minutes to under 2 minutes. The service moves weights directly between GPUs using RDMA via NIXL, bypassing centralized broadcasts. It also reuses JIT-compiled kernel caches across replicas, further reducing latency for inference and RL post-training workflows.

Read more
NVIDIANVIDIAJul 24

NVIDIA Kaggle Team Wins NeuroGolf and KDD Cup Agent Competitions

NVIDIA’s Kaggle Grandmasters team secured 1st place in the 2026 NeuroGolf Championship and 2nd in the KDD Cup 2026 Data Agents Competition. The team used manager agents with specialized workers for NeuroGolf, while their KDD Cup entry utilized a purpose-built harness to narrow the performance gap between a frontier LLM and a smaller open model.

NVIDIANVIDIAJul 23

NVIDIA Shows Hosted RL Customization for Nemotron 3 Nano Models

NVIDIA demonstrated a hosted reinforcement learning workflow using Prime Intellect Lab that improved Nemotron 3 Nano’s math task accuracy from 22% to 91%. The process costs under $5 and produces a downloadable LoRA adapter. This same customization workflow scales to the larger Nemotron 3 Super and Ultra models by updating a single configuration line.

Read more
NVIDIANVIDIAJul 22

NVIDIA Publishes Five Lessons on Improving AI Reasoning from Kaggle Challenge

NVIDIA analyzed results from its Nemotron Model Reasoning Challenge, where 5,000+ participants optimized reasoning workflows using Nemotron-3-Nano-30B models. The findings highlight five practical habits for building reliable reasoning systems, including using verifiable chain-of-thought traces, designing traces to fit token budgets, and separating reusable knowledge from live problem-solving. These techniques were validated on Google Cloud G4 VMs with Blackwell GPUs.

Read more
NVIDIANVIDIAJul 22

NVIDIA Ships 4-Step Cosmos 3 Super Models With 25x Speedup

NVIDIA released 4-step Cosmos 3 Super models that generate images and video up to 25x faster than the original versions. These models now rank first for image-to-video and second for text-to-image on the Artificial Analysis open-weight leaderboards. The new checkpoints are available for use on Hugging Face.

Read more
NVIDIANVIDIAJul 22

Wistron Launches First U.S. Factory to Produce NVIDIA AI Superchips

Wistron opened its first U.S. manufacturing facility in Fort Worth, Texas, to produce the NVIDIA GB300 Grace Blackwell Ultra Superchip. The $700 million plant utilizes digital twin technology to optimize production and will eventually manufacture the Vera Rubin Superchip. This facility expands domestic capacity for assembling and testing advanced AI infrastructure.

Read more
NVIDIANVIDIAJul 21

NVIDIA Nemotron 3 Ultra Achieves Gold Medal Score at IMO 2026

NVIDIA tested its Nemotron 3 Ultra model on the International Mathematical Olympiad 2026 problems under strict competition constraints, including no internet or external tools. The model achieved a score of 30/42, surpassing the official 29-point gold medal threshold. This result demonstrates the model's mathematical reasoning capabilities in a controlled, time-limited environment.

Read more
NVIDIANVIDIAJul 21

NVIDIA GB300 NVL72 Sets World Record for DeepSeek-V3 Pre-Training

NVIDIA GB300 NVL72 achieved a world record of 1,648 TFLOPs per GPU while pre-training the DeepSeek-V3 671B model. This performance represents a 3x increase over the previous-generation GB200 NVL72. The throughput gains result from hardware-software co-design and continuous optimizations across the Megatron-Core, TorchTitan, and JAX frameworks.

Read more
NVIDIANVIDIAJul 20

NVIDIA Releases Cosmos 3 Edge Open World Model for Physical AI

NVIDIA released Cosmos 3 Edge, an open 4-billion-parameter world model designed for on-device physical AI. The model combines autoregressive and diffusion transformer towers to enable real-time scene understanding, prediction, and action generation for robots and autonomous systems. It ranks first on VANTAGE-Bench for vision analytics and includes model weights, post-training recipes, and code on Hugging Face.

Read more
NVIDIANVIDIAJul 20

NVIDIA Agent Toolkit Adds Omniverse Libraries for Physical AI Workflows

NVIDIA updated the Agent Toolkit with new Omniverse libraries, providing AI agents with tools for physical AI simulation. The release includes ovrtx for sensor simulation, ovphysx for GPU-accelerated physics, and CAD-to-SimReady skills for asset validation. These components automate scene inspection and asset preparation within existing 3D applications, with integration blueprints now available on GitHub.

Read more
NVIDIANVIDIAJul 17

NVIDIA Nemotron 3 Embed Models Lead Long-Horizon Memory Benchmark

NVIDIA’s Nemotron 3 Embed 8B and 1B models secured the first and second positions on the Long-Horizon Memory (LMEB) benchmark. The models achieved scores of 64.36 and 61.50 across 22 tasks. This benchmark evaluates retrieval accuracy in long-running conversations and memory-intensive tasks, providing a measure for agentic systems that require context retention across sessions.

Read more
NVIDIANVIDIAJul 17

NVIDIA NeMo AutoModel Adds Distributed Training for Hugging Face Diffusers

NVIDIA updated its NeMo AutoModel library to support Hugging Face Diffusers, enabling distributed fine-tuning for image and video models. The integration provides ready-to-run full and LoRA recipes for models like FLUX.1-dev and Wan 2.1 without requiring checkpoint conversion. Sharding and parallelism configurations now extend directly to Diffusers models, scaling training across multiple GPUs without architecture-specific wiring.

Read more
NVIDIANVIDIAJul 16

NVIDIA Releases Nemotron 3 Embed 8B, Ranking #1 on RTEB

NVIDIA released the Nemotron 3 Embed 8B model, which secured the #1 ranking on the RTEB retrieval accuracy benchmark. The open-weights model supports a 32k context window and multilingual retrieval, aiming to improve agentic context accuracy. NVIDIA also introduced high-efficiency 1B variants, including a Blackwell-optimized NVFP4 version, to support production-scale retrieval and agent memory workflows.

Read more
NVIDIANVIDIAJul 16

NVIDIA and Noetra Build Japan’s First National Physical AI Infrastructure

NVIDIA is partnering with Noetra and Japan’s METI to build the world’s first national AI infrastructure for physical AI. The factory, built on the NVIDIA DSX platform, features 27,500 Rubin GPUs and 13,750 Vera CPUs to deliver 140 megawatts of capacity. It supports Japan’s FRONTia Project, providing foundation models for robotics and industrial applications.

Read more
NVIDIANVIDIAJul 16

NVIDIA DeepStream 9.1 Adds Agentic Skills, 3D Tracking, and Calibration

NVIDIA released DeepStream 9.1, introducing 13 agentic skills to automate video analytics pipeline development. The update features Multi-View 3D Tracking for consistent object IDs across camera networks and AutoMagicCalib for automated camera calibration. It supports JetPack 7.2 for Jetson Orin and Thor edge platforms, with all source code and reference applications now available on GitHub.

Read more
NVIDIANVIDIAJul 15

NVIDIA GEAR Lab Scales Robot Context to 8,000 Timesteps

NVIDIA GEAR Lab introduced RoboTTT, a robot foundation model that scales visuomotor context to 8,000 timesteps—three orders of magnitude beyond current policies—without increasing inference latency. The model uses test-time training to compress history into fast weights, enabling one-shot imitation from human video, on-the-fly self-correction, and improved performance on long-horizon tasks.

Read more
NVIDIANVIDIAJul 14

NVIDIA Automates Cosmos 3 Nano Post-Training to Boost Model Accuracy

NVIDIA automated the post-training of its Cosmos 3 Nano model using TAO agent skills and AutoML. The workflow increased traffic signal detection accuracy on the Woven Traffic Safety dataset from a 54.41% zero-shot baseline to 93.35% in under a day. The process handles data patching, baseline evaluation, and hyperparameter sweeps through natural language prompts.

Read more
NVIDIANVIDIAJul 14

NVIDIA Coding Agent Autonomously Trains Vision Model to 96.9% Accuracy

NVIDIA demonstrated an autonomous coding agent that uses NeMo RL, NeMo Gym, and reusable agent skills to manage reinforcement learning research. The agent built a visual counting environment and trained a Qwen3-VL-2B-Instruct model, increasing accuracy from 25% to 96.9%. It also autonomously proposed a follow-up experiment, while the researcher maintained strategic oversight of the campaign.

Read more
NVIDIANVIDIAJul 10

NVIDIA Research Releases Flex-Forcing for Flexible Video Generation

NVIDIA Research released Flex-Forcing, a video generation method that trains a single model to switch between bidirectional diffusion and autoregressive generation at inference time. This framework allows selection of a generation approach based on specific compute budgets, balancing structural consistency with streaming speed. The project was recognized with a spotlight at ICML 2026.

NVIDIANVIDIAJul 7

NVIDIA Research Introduces MOTIVE for Motion-Centric Video Data Attribution

NVIDIA Research introduced MOTIVE, a motion-centric data attribution framework that identifies training clips influencing temporal dynamics in video generation. By re-weighting gradients toward moving regions, the method enables curation of high-influence subsets. Fine-tuning on these subsets improves motion smoothness and dynamic degree, achieving a 74.1% human preference win rate against the base model.

NVIDIANVIDIAJul 6

ICML Paper Quantifies LLM Memorization Capacity at 3.6 Bits per Parameter

NVIDIA AI highlights an ICML 2026 research paper that estimates the memorization capacity of GPT-style models at approximately 3.6 bits per parameter. This metric distinguishes unintended data memorization from model generalization, offering a quantitative approach to evaluate training data requirements, scaling laws, and privacy risks in large language models.

NVIDIANVIDIAJul 1

NVIDIA Research Releases Nemotron-Labs-TwoTower for 2.42x Faster Text Generation

NVIDIA Research released Nemotron-Labs-TwoTower, a diffusion language model adapted from the 30B-parameter Nemotron-3-Nano-A3B. The architecture splits the model into a frozen context tower and a trainable denoiser tower, enabling parallel token generation. This approach retains 98.7% of the original model’s quality while delivering 2.42× faster wall-clock throughput. Code and weights are available on Hugging Face.

Read more
NVIDIANVIDIAJul 1

NVIDIA Research Introduces Generative Pretrained Controllers for Reusable Motor Control

NVIDIA Research introduced Generative Pretrained Controllers (GPC), a framework that models motor skills as discrete tokens for transformer-based next-token prediction. Trained on 600+ hours of motion data, the controller runs in real-time within physics simulations and achieves a 99.98% success rate in reproducing motion. The framework allows fine-tuning a single pretrained controller to solve various new downstream tasks.

Read more
NVIDIANVIDIAJun 30

NVIDIA GEAR Lab Unveils ASPIRE for Autonomous Robot Skill Discovery

NVIDIA GEAR Lab introduced ASPIRE, a system that autonomously discovers reusable robot skills by writing and refining control programs. It compounds experience into a library of 90+ skills across 150+ tasks, achieving up to a 10x reduction in transfer learning tokens. The system enables sim-to-real and cross-embodiment transfer by shipping code-based know-how rather than neural network weights.

Read more
NVIDIANVIDIAJun 30

NVIDIA TAO 7 Ships with Agent Skills and AutoML Optimization

NVIDIA released TAO 7, adding agent skills that plug into coding agents to improve accuracy. The update includes LLM-guided AutoML that finds hyperparameter configurations up to 2x faster, plus local fine-tuning for HuggingFace CV and VLM models. Data-enhanced fine-tuning also allows agents to identify and fix model failures automatically.

NVIDIANVIDIAJun 28

NVIDIA Unifies Internal AI Model Access with Enterprise Inference Hub

NVIDIA built the Enterprise Inference Hub to manage AI model access for thousands of internal engineers. Using LiteLLM as a central gateway, the platform provides a single API for over 100 models across cloud, open-source, and internal services. The hub now processes trillions of tokens every week, centralizing authentication, monitoring, and cost management for all internal AI applications.

Read more
NVIDIANVIDIAJun 24

NVIDIA NeMo AutoModel Integrates Transformers v5 for Faster MoE Training

NVIDIA NeMo AutoModel now integrates HuggingFace Transformers v5, adding Expert Parallelism, DeepEP, and TransformerEngine kernels. This update delivers 3.4–3.7x higher training throughput and 29–32% lower peak GPU memory usage for Mixture-of-Experts models. The framework maintains API compatibility, enabling these performance gains through a single import line change without requiring additional code rewrites.

Read more
NVIDIANVIDIAJun 23

NVIDIA Adds DFlash Speculative Decoding for 15x Higher Blackwell Inference Throughput

NVIDIA integrated DFlash, an open-source block diffusion model, to accelerate inference on Blackwell GPUs by up to 15x. By drafting entire token blocks in parallel rather than sequentially, DFlash maintains interactivity targets while boosting throughput. The update is available as a drop-in integration for SGLang, TensorRT-LLM, and vLLM, with 20 model checkpoints released on Hugging Face.

Read more

Frequently asked questions

NVIDIA is AI computing company building GPUs, inference hardware, and developer platforms for training and deployment. HeadsUpAI tracks NVIDIA across the AI ecosystem and curates every significant update — the latest being "NVIDIA and USC Unveil HorizonRelight for Consistent Long-Horizon Video Relighting" (September 10, 2026) — so you get the whole story in a 30-second read.

The most recent NVIDIA update is "NVIDIA and USC Unveil HorizonRelight for Consistent Long-Horizon Video Relighting" (September 10, 2026). HeadsUpAI curates every significant NVIDIA release as a 30-second read — what shipped and why it matters.

The latest NVIDIA updates: "NVIDIA and USC Unveil HorizonRelight for Consistent Long-Horizon Video Relighting", "NVIDIA Details Encode-Prefill-Decode Disaggregation for Multimodal Model Serving", "NVIDIA and Palantir Combine AI Stacks for Supply Chain Decision Intelligence", "NVIDIA Partners With Australian Ecosystem to Build 2-Gigawatt AI Factories", and "NVIDIA Nemotron Model Outscores Top Human Coders at IOI 2026". HeadsUpAI has curated 67 NVIDIA updates over the last 90 days, covering company news, research, and product updates — listed newest first, presented straight, no hype, no bias.

NVIDIA is AI computing company building GPUs, inference hardware, and developer platforms for training and deployment. On this page you'll find every significant NVIDIA development HeadsUpAI has tracked recently — company news, research, and product updates — so you can keep up with where NVIDIA is heading without reading a dozen sources.

Continuously. HeadsUpAI adds new NVIDIA updates as they're announced — usually within hours — and the 67 updates currently shown cover the past 90 days, newest first.