Inference is moving to CPUs: one filed sentence, zero 10-Q mentions, and a $5.0 billion stake
Every figure below was read from the filed document on sec.gov and is cited with its accession.
Word counts — "the word appears zero times" — come from a case-insensitive search for the stem
inferenc over the raw filing HTML, so a hit buried in markup would still be caught. Phrase counts
across all filers come from EDGAR full-text search, which indexes filing documents from 2001 onward;
it counts documents, not filings, and covers neither earnings-call audio nor image-based exhibits.
13F figures were read out of each filing's information table XML and reconciled to the cover-page
total. Arithmetic we did on filed values is labelled as ours. Research, not investment advice.
There is a trade a lot of people are in right now with a one-line summary: the next leg of AI compute is inference, inference is cheaper and more general than training, and that pushes work back toward CPUs — so the CPU names are the ones to hold.
It is a coherent story. We wanted to know how much of it is written down somewhere a company can be held to it. So we read the filings — Intel's, NVIDIA's, AMD's, Arm's, then the hyperscalers' and the server builders'. The answer is more interesting than "true" or "hype", because the filings and the headlines are not making the same claim.
Start with the sentence the thesis rests on
One filed sentence anchors this story, in Exhibit 99.1 to Intel's Form 8-K, accession
0000050863-25-000155, filed September 18, 2025:
For data centers, Intel will build NVIDIA-custom x86 CPUs that NVIDIA will integrate into its AI infrastructure platforms and offer to the market.
That is the entire data-center commitment. Alongside it, in Item 3.02, is the money:
NVIDIA agreed to purchase 214,776,632 shares of the Company's common stock … at $23.28 per share, representing an aggregate purchase price in cash of $5.0 billion.
Now notice what is absent. The word "inference" does not appear in that 8-K or its press release — the count is zero. The release frames the deal as "AI infrastructure platforms"; the exhibit's headline is "Intel to design and manufacture custom data center and client CPUs with NVIDIA NVLink". Neither CEO quote uses the word. The sentence carries no volume, no revenue and no date, and the exhibit's own forward-looking block flags "the risk that the anticipated benefits of the collaboration agreement may not materialize."
Two details got compressed out of the coverage: the purchase was signed in September 2025 but
closed on December 26, 2025 (Intel 10-K 0000050863-26-000011), and the other half of the deal
is PCs — "x86 system-on-chips (SOCs) that integrate NVIDIA RTX GPU chiplets."
The phrase does not exist in EDGAR
So we asked the simplest question available: how often do companies file the claim?
| Exact phrase | 2026 (Jan 1 – Sep 22) | Ever (2001 – present) |
|---|---|---|
"CPU inference" | 0 | 0 |
"inference on CPUs" | 0 | 0 |
"CPU-based inference" | 0 | 0 |
"CPUs for inference" | 0 | — |
"general-purpose CPUs for inference" | 0 | — |
Zeros invite suspicion, so here are controls on the same endpoint, against the same 265-day window a year earlier:
| Control phrase | 2026 | 2025 | Change |
|---|---|---|---|
"inference workloads" | 340 | 49 | 6.9x |
"inference" | 5,342 | 4,359 | +22.6% |
"Trainium" | 16 | 4 | 4.0x |
"custom silicon" | 112 | 44 | 2.5x |
"Graviton" | 25 | 24 | +4.2% |
The index is live, and the language is genuinely moving: "inference workloads" went from 49
documents to 340. That is the number to watch, not the bare word — "inference" alone is
contaminated, its top 2026 filers being Goldman Sachs structured-note prospectuses, a
mineral-resources issuer, and the ordinary legal phrase "adverse inference". We show it only as a
control.
What survives is a precise negative: the CPU-inference phrasing has never been filed by anyone, in any form, since the corpus begins. That is not proof no CPU runs inference — plainly many do. It is evidence about something narrower: no company has put its name to that claim in a document it can be held to. Worth noting too that none of the five large hyperscalers appears in the "inference workloads" results at all; scoped searches return zero for each. The phrase belongs to the tier below them.
Intel's best data-center quarter in years never says it either
Intel's 10-Q, accession 0000050863-26-000157, filed July 24, 2026, quarter ended June 27, 2026
covers the quarter people point at. The word counts over the raw document:
| Term | Occurrences |
|---|---|
inference / inferencing | 0 |
Xeon | 0 |
NVIDIA | 0 |
AI workload | 0 |
host CPU | 0 |
The prior quarter (0000050863-26-000079) is identical — all zero. Intel's data-center segment,
filed as "Data Center and Artificial Intelligence" (DCAI), did $6,262 million against $3,939
million, up 59%, at a 40% operating margin. The filing explains that without reference to AI:
DCAI revenue increased $2.3 billion from Q2 2025 … primarily driven by higher server revenue … due to ASP increases of 48% and 38%, respectively. The majority of the increase in server ASPs was driven by a higher mix of premium products sold … with demand-based pricing actions contributing to a lesser extent, in part to offset higher input costs. Server volume increased 9% compared to Q2 2025 … primarily driven by higher hyperscaler demand. Market demand exceeded our available product supply … due to internal supply constraints.
Read as a decomposition, that is the shape of a commodity squeeze: price up 48%, volume up 9%, supply short. The prior quarter said the same with smaller numbers — ASPs +27%, volume down 5%. The FY2025 10-K said it with ASPs going the other way entirely: "Server ASPs decreased by 4% from 2024 … an increase in server volume of 9%."
The 10-K adds a detail that complicates any "new AI CPU" reading: a majority of 2025 data-center revenue came from 3rd, 4th and 5th Gen Xeon parts built on Intel 7, a node Intel has shipped for years. (Our inference, labelled as ours: old silicon at high prices fits tight supply and premium mix more neatly than a new class of workload. The filing states the mix; the reading is ours.)
Where Intel does say "inference" — and what it names
Intel uses the word eight times in its FY2025 10-K. Every one is in Item 1, the business and strategy section. Zero in MD&A, zero in Risk Factors — and that placement is the point, since Item 1 is where a company states ambitions and MD&A is where it explains what happened.
Inference AI, agentic AI and physical AI workloads are rapidly emerging areas that we expect may ultimately represent larger market opportunities than the generative AI workloads … We also aim to participate in these emerging areas through our continued development of our Xeon, AI PCs, Arc GPUs and our open software stack, as well as by developing successive generations of inference-optimized GPUs on a targeted annual cadence …
Two things. The hedge is doubled — "we expect it may ultimately". And the product Intel names as its inference vehicle is a GPU: Crescent Island, "our first GPU … optimized performance for AI inference workloads."
Intel's own description of where a CPU sits is the accurate one, and it is not "instead of":
We pair our x86 CPUs with GPUs, IPUs, NPUs and other accelerators … This heterogeneous approach enables us to deliver compute platforms that match the specific requirements of inference, training and orchestration tasks.
NVIDIA: one inference sentence, and "CPU" appears zero times
Take the company whose revenue is the inference boom. NVIDIA's 10-Q 0001045810-26-000075,
filed August 26, 2026, quarter ended July 26, 2026 reports Data Center revenue of $89,023 million
against $41,096 million, up 117%. The stem inferenc appears exactly once:
We believe AI clouds and AI model makers have significant demand for training and inference compute and currently lack the ability to secure long-term infrastructure contracts and investment-grade financing capacity to secure the AI infrastructure necessary to grow.
A sentence about financing, welding inference to training rather than separating them. In the same
document: Intel 0, x86 0, Grace 0, and the string CPU 0. The company that
bought $5.0 billion of Intel stock mentions neither Intel nor CPUs in its latest quarterly report.
In NVIDIA's FY2026 10-K (0001045810-26-000021), the only CPU attached to inference is its own:
In 2024, we launched the NVIDIA Blackwell architecture – connecting 36 Grace CPUs and 72 Blackwell GPUs in a data center scale, liquid-cooled design – for real-time trillion-parameter inference and training.
(Our arithmetic: one CPU per two GPUs, described as a component of a GPU rack.) Intel appears six times in that 10-K, every one in a competitor list.
AMD files one number where the story needs two
AMD is the cleanest test, because it sells both sides. Data Center revenue was $6.7 billion, up
107% from $3.2 billion (10-Q 0000002488-26-000123; results exhibit 0000002488-26-000121, dated
August 4, 2026). The attribution:
The increase in both periods was primarily driven by strong demand for our AMD EPYC processors and AMD Instinct MI350 Series GPUs.
CPUs and GPUs, one line, no split. AMD does not disaggregate CPU from GPU revenue in any filing we read, so the single most important number for this thesis — how much of that $6.7 billion is EPYC — is not filed and cannot be computed from the documents.
The word appears zero times in AMD's 10-Q, and four times in the results exhibit, always attached to a GPU or a rack. This one is worth reading closely:
Expanded collaboration with Microsoft to deploy AMD Helios racks at scale on Azure to power frontier model inference. Azure will also add two new AMD EPYC CPU-powered VM series …
In one filed sentence, inference is what the Helios racks do; the EPYC CPUs are a different product
in the next clause. AMD's FY2025 10-K taxonomy draws the same line in consecutive sentences — EPYC
CPUs "designed for high-performance computing, enterprise IT, supercomputing, and large data
centers", Instinct GPUs "designed for AI training, inference and exascale-class scientific
computing." Across all seven inference occurrences in that 10-K and all four in the exhibit, not
one attaches the word to EPYC or any AMD CPU.
Arm files the clearest CPU-and-AI sentences — read them precisely
The best CPU evidence is not in an annual report. It is in Arm's Form 6-K, Exhibit 99.2, accession
0001973239-26-000113, filed July 29, 2026, the shareholder letter for the quarter ended June 30,
2026: "data center royalties more than doubled year over year", with Neoverse shipments past
1.5 billion cores, "the most recent 500 million shipping in just nine months, where the first
1 billion took 6 years." Then the roll-call:
NVIDIA announced that Vera has entered full production. Built on Arm, Vera delivers up to 50% higher CPU performance and 2x greater energy efficiency than comparable x86 systems and will serve as the CPU foundation for NVIDIA's next-generation AI infrastructure … Google has stated that its Arm-based Axion CPU is a core component of its AI infrastructure strategy, highlighting its role as the host CPU for Google's latest TPU AI systems. AWS also expanded momentum behind its Arm-based Graviton platform, announcing a multi-year agreement with Meta to deploy tens of millions of Graviton5 cores to power agentic AI workloads. Microsoft expanded Azure Cobalt 200 virtual machines built on Arm Neoverse CSS. And Qualcomm has announced plans to enter the AI data center CPU market with its Arm-based Dragonfly C1000.
This is simultaneously the strongest CPU evidence in the set and a check on the strong version of the thesis. Axion's filed role is "the host CPU for Google's latest TPU AI systems" — the CPU hosts the accelerator, it does not replace it. Graviton5 cores "power agentic AI workloads", which is not the same claim as running inference. Vera is the CPU "foundation" for NVIDIA's infrastructure. And the letter containing all of it uses the word "inference" zero times.
Arm's 20-F for the year ended March 31, 2026 uses the word exactly once, as a power problem: "Data centers must cope with rising AI training and inference needs while limited by the power available from the grid." Royalty revenue was $2,613 million against $2,168 million, up 21%. The 20-F discloses no data-center share of royalty and never names Graviton, Axion, Cobalt or Grace — zero occurrences of each. The detail lives in the quarterly 6-K; the annual report is generic.
The buyers separate the two jobs explicitly
Four of the five large hyperscalers barely acknowledge their own chips in a periodic report.
Alphabet has never written "Axion" — its own Arm server CPU — in any SEC filing; zero hits, all
time, while it names TPUs freely and now books revenue from selling TPU systems (10-Q
0001652044-26-000071). Microsoft has not named Maia in a filing since FY2024 and has never
named Cobalt outside conflict-minerals filings, where the word means the metal. Meta has never used
the word "inference" in any SEC filing — zero hits, all forms, all time.
Amazon is the exception, and instructive. Graviton, Trainium and custom silicon appear zero
times in its latest 10-Q (0001018724-26-000026) and 10-K; the entire story is filed on 8-K
exhibits. From the 2026 shareholder letter (0001104659-26-041034, April 9, 2026): "virtually all of
the workloads ran on Intel chips until we invented Graviton in 2018"; Graviton "is now used
expansively by 98% of the top 1,000 EC2 customers"; and the chips business — Graviton, Trainium and
Nitro combined — is "now over $20 billion, and growing triple digit percentages YoY." That is the
largest filed number in this story, and Amazon does not split CPU from accelerator either.
Then it assigns the two jobs to two different chips, in writing: "lower-cost inference (on our custom silicon, Trainium)", and "Amazon Bedrock, AWS's primary (and very fast-growing) inference service, runs most of its inference on Trainium."
Trainium is an accelerator. The company with the deepest custom-silicon programme in the industry
routes its inference to the accelerator and says so repeatedly. So what is the CPU for? The Q1 2026
earnings 8-K (0001018724-26-000012) answers it in the best filed sentence on CPUs and AI anywhere:
Signed an agreement with Meta … to deploy tens of millions of AWS Graviton cores to power CPU-intensive workloads behind its agentic AI efforts, including real-time reasoning, code generation, and multi-step agent workflows.
Not "inference" — "CPU-intensive workloads behind its agentic AI efforts." Behind. The CPUs run retrieval, orchestration, tool calls and code execution around the model; the model runs elsewhere. Arm describes the same agreement from the other side of the contract — "tens of millions of Graviton5 cores to power agentic AI workloads" — and the two filings corroborate each other. (Our inference, labelled as ours: two filers describing one agreement in near-identical words is strong evidence it exists and is large. Neither gives a dollar value, so its size is not measurable.)
The one filing that moved toward the thesis — and it is a wording change
Dell's FY2026 10-K (0001571996-26-000008, filed March 16, 2026) described its traditional
server line as "supporting a wide range of general-purpose and mission-critical workloads." Three
months later, in the Q1 FY2027 10-Q (0001571996-26-000030, filed June 9, 2026) — and again in
the Q2 FY2027 10-Q (0001571996-26-000046, filed September 8, 2026) — the same bullet had grown
a clause:
… supporting a wide range of general-purpose and mission-critical workloads, including certain AI-related workloads such as inferencing.
That is the closest thing to the thesis in filed form anywhere in this research: inferencing moved, in writing, into the description of a general-purpose server line, on a date you can point at.
Then look at the numbers underneath it. Dell's Infrastructure Solutions Group, quarter ended July 31, 2026:
| Line, as filed ($M) | Q2 FY2027 | Q2 FY2026 | Change |
|---|---|---|---|
| AI-optimized servers | 16,401 | 8,208 | +100% |
| Traditional servers and networking | 10,531 | 4,736 | +122% |
| Storage | 4,850 | 3,856 | +26% |
| Total ISG net revenue | 31,782 | 16,800 | +89% |
Traditional servers grew faster than AI-optimized ones. But Dell files a driver for each, and they differ: AI-optimized revenue rose "primarily driven by an increase in units sold", traditional servers "primarily driven by an increase in the average selling price." Units on one line, price on the other — and Dell says where the price came from: "current limitations in capacity from memory manufacturers … substantial inflation in memory component costs." A year earlier the FY2026 10-K reported traditional servers up 9% with units falling.
HPE says it with less ambiguity. Its Q3 FY2026 10-Q (0001645590-26-000080, September 3, 2026)
reports Server revenue of $6,766 million against $5,000 million, up 35.3%, "predominantly due to
an increase in the average selling price … primarily driven by commodity price increases, especially
memory and SSDs." A general-purpose server line growing a third, attributed entirely to price,
attributed entirely to memory. HPE's 10-Q uses "inference" zero times.
Then the hardest number in the set, from a place nobody looks. Super Micro runs a CPU-based revenue
KPI in its executive incentive plan, disclosed in its FY2026 10-K (0001375365-26-000022, filed
August 31, 2026). It pays zero if CPU-based revenue as a share of total revenue lands below the prior
year. The filed result for "CPU based Revenue" was 0%, footnoted: "CPU based revenue in certain
processor categories decreased compared with fiscal year 2025." Super Micro's net sales rose 77.8% to
$39,063.1 million that year. Its CPU-based revenue share fell, and it paid nothing on that measure.
IBM adds the budget picture (0000051143-26-000078, July 23, 2026): "Many clients redirected
spending toward servers, storage, and memory purchases to secure supply-constrained infrastructure
ahead of expected price increases," with Distributed Infrastructure revenue up 37 percent, "our
strongest quarter on record." Enterprises are buying general-purpose infrastructure hard, and IBM's
filed reason is securing supply ahead of price rises — not a new class of workload.
The exception, and it is a real one
Fairness demands the counterexample, and there is exactly one. IBM does attach inferencing directly
to a CPU — its own — in its FY2025 10-K (0000051143-26-000010, filed February 24, 2026):
IBM Z: the premier transaction processing platform … Powered by the IBM Telum processor — which delivers integrated, real-time inferencing and advanced security features — the platform supports emerging generative and multi-model AI workloads.
The same filing attributes revenue to it: "IBM Z revenue increased 51.7 percent … Clients are investing in z17 for its differentiated capabilities, real-time AI inferencing, quantum-safe security …" And the July 2026 10-Q repeats it for a new LinuxONE system offering "the security, resiliency and real time inferencing the IBM Z platform can deliver."
Note two things. IBM says "inferencing", not "inference", which is why a naive keyword search misses
it entirely — our counts use the stem inferenc precisely to catch this. And the number moved
against it: IBM Z revenue fell 42.0% year over year in Q2 2026, per the same 10-Q. One filer has
written the on-thesis sentence down, about a mainframe processor, in a line that is currently
shrinking.
The memory side, briefly
Inference demand lands on memory wherever it runs, so one paragraph — with the caveat that
Micron's 10-Q (0000723125-26-000015, quarter ended May 28, 2026) uses the word "inference" zero
times. It says "AI".
Revenue was $41,456 million against $9,301 million, DRAM $31,328 million against $7,071 million (Note 14), income before taxes $33,212 million against $2,113 million. The decomposition is Intel's shape, more extreme: "Sales of DRAM products increased 343%, primarily due to a low-260% range increase in average selling prices and a low-20% range increase in bit shipments." The filing now also carries committed volume — remaining performance obligations of "approximately $5 billion, of which $422 million has been recognized as contract liabilities", against "not material" a year earlier, under take-or-pay agreements "with binding commitments for specific volumes". Micron files its own hedge: "Although AI is a relatively new demand driver for our products, it is evolving rapidly, and the expected timing and amount of demand related to AI can change significantly."
We covered who absorbs those prices downstream in the memory-cost squeeze on server OEMs, and are not repeating it here.
Who actually owns the CPU names
One filed angle with no interpretation in it, read out of the information tables for the quarter
ended June 30, 2026. NVIDIA's Form 13F-HR (0001045810-26-000065, filed August 14, 2026) reports
eight positions totalling $63,439,974,569, of which Intel is 214,776,632 shares at
$29,989,261,126 — 47.3% of the book, larger than its SpaceX line, and matching the 8-K share count
exactly. SoftBank Group's 13F (0001062993-26-004387) carries Intel at $12,141,739,167 on
86,956,522 shares — 66.8% of its entire $18,174,540,337 book. Both mark Intel at exactly $139.63, a
clean cross-check that the two XMLs agree.
(Our arithmetic: $139.63 against the $23.28 purchase price is roughly 6x on $5.0 billion of cash; the two filers together hold 301,733,154 shares against the 5,043 million Intel's balance sheet shows outstanding at June 27, 2026 — 5.98% of Intel.)
The edge runs one way. Intel's own 13F (0000050863-26-000176) is three lines totalling $620,621,529
— Joby Aviation and Mobileye — and AMD's (0001193125-26-352454) is six lines totalling
$1,311,750,821. Neither holds the other, and neither holds NVIDIA.
What the filings separate, and what the headlines merge
Price versus volume. Intel's data-center growth is filed as ASP +48%, volume +9%. Dell's traditional servers are filed as price; its AI servers as units. HPE's 35% is filed as price, from memory. Micron's is ASP +260% range against bits +20% range. MD&A gives the drivers because the SEC asks for them. "Data-center revenue up 59%" carries none of it, and the two components do not behave alike.
Committed versus indicative. A take-or-pay obligation with a binding volume is a different object from a collaboration announcement. Micron files a number for the first. Intel's NVIDIA collaboration has no filed volume, revenue or date — a sentence and a hedge. Both are real; one is measurable.
The chip versus the system. This is what the story turns on. Almost every filing that puts a CPU near an AI workload describes it as part of a system: Grace at 36 per 72 Blackwell GPUs; Axion as the host CPU for Google's TPUs; Graviton for the work behind agentic AI while Bedrock's inference runs on Trainium; EPYC beside Instinct in one undisaggregated segment. The single exception is IBM's Telum, a mainframe processor, in a product line down 42% year over year.
So the honest reading depends on which claim is being made. If it is "CPU attach grows as inference fleets grow", the filings are consistent with it and do not quantify it — and Dell's new clause is the first wording to lean that way. If it is "inference is migrating off GPUs onto general-purpose CPUs", one filer has said something adjacent about a shrinking mainframe line, nobody has said it about volume servers, and Super Micro's own CPU-revenue KPI went the other way. What no filing supports in either case is a claim about market share, because the number that would settle it — a CPU-versus-accelerator revenue split — is not disclosed by Intel, AMD, Amazon or Dell.
What we are not doing is calling the cycle. Intel says supply constraints "persist into next year"; Micron says AI demand "can change significantly"; AMD's outlook is explicitly forward-looking. Those are the companies telling you the range is wide.
Run it yourself
The fastest first step is the one that produced the tables above: search the phrase, not the theme.
EDGAR full-text search
takes an exact phrase, a form type and a date range, and a zero result is information. Run
"inference workloads" against "CPU inference" and watch the gap.
Then read the decomposition rather than the headline. Intel's financials and AMD's give you the standardized statements, /filings/INTC takes you to the documents, and /holders/INTC shows who reports owning it — with /holders/NVDA/owns turning the data around to show NVIDIA's own disclosed book, Intel line included. Every number above traces back from those pages to an accession.
If you want the filings inside an agent instead, the same search and statement tools are on the EvidInvest MCP server. It is free to start, credit packs begin at $10, and there is no subscription.
The reason to do this before the next quarter is that the interesting change here would be a wording change — Dell's was. The moment Intel's MD&A, not its business section, attributes revenue to inference, or AMD splits EPYC out of Data Center, the story stops being an inference and becomes a number. Our Thesis Monitor reads each new filing against what you already believed, which is the part a quarterly headline cannot do on its own.
Research, not investment advice.
Working with this data from an AI agent? The EvidInvest MCP server gives Claude, Cursor, and any MCP client access to 46 financial data, valuation, and SEC intelligence tools.
Fair Value Weekly
Get DCF breakdowns, fair value updates, and portfolio ideas for serious investors. No spam, no paywalled teasers.