Forget the marketing for a second and read what NVIDIA had to tell the SEC. Its FY2026 annual report describes accelerated computing for workloads "such as artificial intelligence, or AI, model training and inference, data analytics, scientific computing, robotics, and 3D graphics." Two of those words — training and inference — are the whole game, and the filing lists them separately on purpose.

The filing is even more explicit a few pages on, where it locates these two jobs at the literal foundation of the company's hardware. It states plainly what the chips are for.

At the foundation of the NVIDIA accelerated computing platform are our GPUs, which excel at parallel workloads such as the training and inferencing of neural networks.— NVIDIA Corp Form 10-K (FY2026) source

The way this actually works: training is construction. You show a model enormous amounts of data and repeatedly nudge its internal parameters until its outputs are good. It is intermittent and brutally compute-hungry — you do it to build a model, not every time you use one. The sec.gov filing (surfaced via SEC filings) groups it with the other heavy scientific jobs because, mechanically, that is what it is: a large, parallel optimization problem run across enormous grids of numbers.

Why the filing keeps the two verbs apart

Inference is the opposite shape. It is running the trained model forward to produce one answer — a sentence, an image, a classification. Each act is cheap relative to training, but it happens continuously and scales with how many people use the product. Over a model's life, inference is the cost that never stops. The 10-K's own phrasing — "training and inferencing of neural networks" — is careful to name both as the parallel workloads its GPUs are built for, rather than collapsing them into a single "AI" label, and that precision is the point worth borrowing.

The filing also makes clear who is doing each job, and the language quietly spans both phases. It describes how "AI model makers use our infrastructure and software hosted at CSPs to develop, build and run AI models, product offerings, and services" — develop and build is training; run is inference. The same document notes that enterprises and startups use the platforms "to build new generative and agentic AI-enabled products and services," and that the broader software stack — the libraries, SDKs, and APIs the filing enumerates — exists to "accelerate the performance and facilitate the deployment" of these workloads. Deployment is the tell: a model that has finished training still has to be served, and serving is inference at scale.

Why the distinction matters to a non-specialist

Because the AI-spend debate hinges on it. A company can justify a big one-time training bill if the resulting model earns its keep in inference. When you hear arguments about whether AI capex "pays off," what's really being argued is the training-to-inference ratio over a model's lifetime — how much it cost to build something versus how much value it returns every time it answers. The filing groups training with "data analytics, scientific computing, robotics, and 3D graphics" precisely because, from the silicon's point of view, all of these are the same kind of dense parallel math; what changes is whether you are paying that cost once to build something or continuously to use it.

The disclosure context underlines the stakes without overstating them. The same FY2026 report describes the company's newest data-center-scale offerings as built with "extreme co-design," where the GPU is paired with CPUs, NVLink switches, DPUs, NICs, and scale-out networking into "data center-scale computing solutions." That architecture exists because training the largest models no longer fits on one chip — it is spread across whole racks — while inference for those same models has to be cheap and fast enough to serve millions of requests. Both phases pull on the same hardware roadmap from opposite ends, which is why a single annual report has to name them both.

Here's the honest gloss: the 10-K lists these workloads to describe what the products address, not to promise demand for either. That separation — what's disclosed versus what's inferred — is exactly the discipline a careful reader should keep. The filing tells you the chips do training and inference, and that the GPUs "excel at" both; it does not tell you who keeps buying them, or in what mix. Those are questions for the market, not the disclosure.

It is worth noticing how the filing frames the customers on each side of the line, because that is where training and inference stop being abstractions. The report describes a spectrum of buyers using the platform: cloud service providers that host the infrastructure, "AI model makers" who develop and build models on it, and the enterprises and startups who then "build new generative and agentic AI-enabled products and services." The model maker lives mostly in the training world; the enterprise shipping a chatbot or an autonomous-driving stack lives mostly in the inference world, running a finished model over and over. The same accelerated-computing platform serves both, which is why a single annual report has to describe a stack of "libraries, software development kits, or SDKs, and application programming interfaces, or APIs" general enough to cover building a model and deploying it.

If you remember one thing, make it the shape: training builds, inference serves. Training is the intermittent, capital-heavy job of fitting a model's parameters to data; inference is the continuous, per-request job of running that finished model forward. Everything else in the AI-hardware conversation — capex debates, "GPU shortages," serving-cost claims — is a footnote on those two verbs, both of which NVIDIA put in writing in its annual report and which you can read yourself in the sec.gov filing.