NVIDIA opens Nemotron datasets: 10 trillion tokens for training AI agents
On July 8, 2026, NVIDIA released the open Nemotron datasets for AI agents: more than 10 trillion pre-training tokens and millions of labeled examples across eight domains, including tool-calling failures. Nemotron-Personas synthetic personas cover 2.4 billion people from 10 countries. The company also launched the interactive Prompt Atlas to explore the corpus.
AI-processed from Hugging Face Blog; edited by Hamidun News
On July 8, 2026, NVIDIA published on Hugging Face a research post about open data for AI agents, while simultaneously releasing Nemotron datasets containing over 10 trillion pre-training tokens and millions of post-training examples. This is one of the largest public releases of training data specifically oriented toward agent systems.
Why model weights alone are insufficient
AI agent behavior is determined not only by neural network weights—but also by the specific data it was trained on. Standard benchmarks and general text corpora do not reflect the real complexity of agent tasks: multi-step action chains, tool use, failures and recovery from errors.
Nemotron datasets cover eight key domains:
- Software engineering traces
- Tool-use failures
- Multi-step reasoning
- Search and information retrieval
- Security and behavior control
- User simulation
- Workflow execution
- Answer verification
The pre-training section exceeds 10 trillion tokens; the post-training section contains millions of labeled examples across multiple domains. Particularly valuable is the inclusion of failure data—erroneous tool calls and failed action chains. Most open datasets capture only successful scenarios; failure data allows agents to learn recovery rather than only achieving results on the first try.
How synthetic data protects confidentiality
One of the main obstacles in developing corporate AI agents—the inability to openly share real workflows. Internal pipelines, client data, and business logic are closed by definition: companies cannot publish what constitutes their competitive advantage.
Synthetic data solves this problem: organizations generate useful training signals from real processes without revealing the processes themselves. This opens the path to participation in open data ecosystems without risking leakage of commercial information.
NVIDIA emphasizes that this approach requires transparency: one must document what was generated, on what basis, and what passed expert review. Synthetic data is not a universal solution—its quality directly depends on the correctness of generator models and validation processes.
"Synthetic data enables participation in open data ecosystems without revealing confidential workflows or client data"—the
Nemotron team.
Prompt Atlas and 2.4 billion synthetic personas
For researchers, NVIDIA created Nemotron Post-Training v3 Prompt Atlas—an interactive visual tool for exploring millions of examples. Data is organized by dataset, domain, and tool-calling patterns, making the corpus not just a downloadable archive but a space for analyzing distributions and patterns.
Nemotron-Personas deserves separate attention—a set of synthetic personas that as of July 8, 2026 covers over 2.4 billion people from ten countries. The personas are "locally grounded": they reflect regional specificity, cultural patterns, and population demographic diversity. This enables training agents that understand local context in requests—not only universal technical English, but communication patterns characteristic of specific regions.
What this means
The release of Nemotron datasets sets a new standard for openness in AI agent development. Inclusion of tool failure data, synthetic personas, and an interactive atlas suggests NVIDIA views agents not as upgraded chatbots but as systems operating in the real world with real users and real failures. Teams building specialized agents now have a ready foundation for fine-tuning—without needing to collect annotations from scratch.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.