AI PC NPUs Explained: What TOPS Ratings Actually Mean for Laptop Performance
Across flagship AI laptops in 2026, TOPS has become the headline number attached to the NPU. Here’s what that number measures, what Microsoft’s Copilot+ certification actually requires, and where the memory bandwidth feeding the NPU — not the compute rating on the box — is the real constraint on performance.
Table of Contents
Quick answer: An NPU’s TOPS rating is a peak theoretical figure, typically measured at a specific numerical precision (usually INT8) under ideal conditions. Microsoft requires 40+ TOPS from the NPU alone for Copilot+ PC certification. Above that floor, TOPS figures from different vendors aren’t directly comparable — they vary by SKU within the same product family, by precision basis, and by whether a vendor is quoting the NPU alone or a combined CPU+GPU+NPU platform total. The bottleneck that actually determines real-world performance is memory bandwidth to the NPU, not the compute rating itself.
When Intel launched its Lunar Lake platform, the headline number was the NPU: 48 TOPS, more than four times what the previous Meteor Lake generation offered. Intel also made two related but distinct architectural decisions that drew far less attention. Inside the NPU itself, Intel is reported to have doubled memory bandwidth and doubled DMA capacity generation-over-generation, specifically to reduce bottlenecks on heavier workloads like large language models. Separately, at the system level, Intel moved up to 32GB of LPDDR5X memory from the motherboard onto the processor package.
That’s the kind of change that shows up in a bill of materials, not a marketing slide. Moving memory onto the package consumes package substrate area, adds assembly complexity, and adds package-level BOM cost that motherboard-mounted memory wouldn’t — while eliminating user-upgradeable memory. Intel spent that budget anyway. The NPU-internal bandwidth and DMA work is Intel’s stated engineering rationale for the workload-stress problem specifically; the on-package memory move is best read as the same instinct applied at the system level, even though Intel’s own public framing for on-package memory emphasizes power draw and motherboard footprint rather than NPU feeding specifically — the connection between the two decisions is this article’s analysis, not a claim Intel has made explicitly.
That’s the part of the “AI PC” story that never makes it onto a spec sheet. The number that does — TOPS — has become the entire conversation, and it’s the wrong one to have in isolation. This piece walks through what an NPU actually does, where Microsoft’s 40 TOPS Copilot+ threshold comes from, how current vendor ratings compare, and why the real bottleneck sits somewhere the spec sheet doesn’t show.
What an NPU Actually Does
An NPU is architected primarily around dense tensor operations — the matrix multiply-accumulate math that neural-network inference runs, over and over, at low numerical precision — surrounded by dedicated data-movement, prefetch, and preprocessing hardware built to keep that math fed. It isn’t a single monolithic block for “one operation”; Intel’s own Lunar Lake NPU disclosure, for instance, describes multiple compute tiles alongside DSP cores and DMA engines, not one undifferentiated array. A CPU prioritizes programmability and general-purpose flexibility. A GPU prioritizes massively parallel throughput, and can be highly power-efficient for workloads suited to it. An NPU narrows further still, dedicating more of its silicon to tensor execution and data movement in exchange for a much narrower supported workload set — continuous, low-power inference without spinning up a fan or pulling down a battery, which is exactly the profile Windows needs for background tasks like live captioning, semantic indexing, and camera segmentation that now run for the length of a session, not the length of a click.
That’s the actual design trade, and it means the CPU, GPU, and NPU aren’t racing on one performance axis. They’re partitioned by workload and power budget. Coverage that reduces “AI performance” to a single number owned by a single block has already misread the architecture.
That specialization has a second-order consequence worth flagging for anyone speccing around this hardware: a CPU’s general-purpose design lets old software keep running acceptably on new silicon, and vice versa, for years. An NPU’s efficiency gains are tied much more tightly to a specific numerical precision and dataflow pattern — which is part of why a chip’s rated TOPS can undersell or oversell what a given model actually achieves on it, depending on how well that model’s format matches the hardware’s fixed execution path. The TOPS number describes the silicon in isolation; it says very little about the fit between that silicon and the specific software running on it.
Where Microsoft’s 40 TOPS Copilot+ Requirement Comes From
The number every buyer now sees is 40 TOPS — trillion operations per second — because that’s the floor Microsoft set for its Copilot+ PC designation. Microsoft’s own developer documentation describes a Copilot+ PC as hardware “powered by a high-performance Neural Processing Unit (NPU) … that can perform more than 40 trillion operations per second,” alongside a minimum of 16GB of RAM and 256GB of storage, and names the qualifying silicon explicitly: AMD’s Ryzen AI 300 series, Intel’s Core Ultra 200V series, Qualcomm’s Snapdragon X series.
That threshold gates specific functionality, not a badge for its own sake. Microsoft’s own framing is narrower than a blanket cutoff, though: its consumer guidance describes 40+ TOPS as what you need for “the best” on-device AI experiences, and it’s the named Copilot+-exclusive features — Recall, Live Captions with real-time translation, Windows Studio Effects, Cocreator — that are documented as requiring the full hardware tier. Below the line, those specific experiences aren’t rated to run; that’s different from claiming no on-device AI of any kind works on lesser hardware. In procurement terms, that’s a defensible engineering floor for that named feature set, not a badge chosen to look good in a press release.
Above the line is where the number stops being a clean signal.
Current Laptop NPU TOPS Ratings by Vendor
TOPS ratings across the four major laptop-silicon vendors have moved fast, and they don’t mean the same thing from one vendor to the next. Based on vendor-published specifications:
| Vendor / Platform | NPU | Published Peak TOPS | Precision Basis | Copilot+ Eligible |
|---|---|---|---|---|
| Qualcomm Snapdragon X Elite / X Plus | Hexagon NPU | 45 TOPS (all SKUs) | Not disclosed in vendor product briefs reviewed | Yes |
| Qualcomm Snapdragon X2 Elite / X2 Elite Extreme | Hexagon NPU | 80–85 TOPS depending on SKU (lineup expanded after launch; some later SKUs reach 85) | INT8 | Yes |
| Intel Core Ultra 200V (“Lunar Lake”) | NPU 4 | Up to 48 TOPS (varies by SKU: 40–48) | INT8 (FP16 also supported) | Yes |
| Intel Core Ultra Series 3 (“Panther Lake”) | NPU 5 | Up to 50 TOPS | INT8 | Yes |
| AMD Ryzen AI 300 (“Strix Point”) | XDNA 2 | Up to 50 TOPS (55 on the Ryzen AI 9 HX PRO 375 specifically) | Not disclosed in reviewed sources | Yes |
| AMD Ryzen AI 400 (“Gorgon Point”) — mobile | XDNA 2 | SKU-dependent; up to 60 TOPS on the flagship, 50 TOPS on lower-tier mobile parts (confirmed for the Ryzen AI 5 435) | INT8 (confirmed for the AI 5 435 and AI 5 PRO 435) | Yes |
| AMD Ryzen AI 400 — desktop | XDNA 2 | Up to 50 TOPS | Not disclosed in reviewed sources | Yes (first desktop Copilot+ chips) |
| Apple M4 | Neural Engine | 38 TOPS | Reported as INT8 by industry press; not stated as such in Apple’s own release | N/A — not a Windows/Copilot+ platform |
| Apple M5 | Neural Engine | Not disclosed | N/A | N/A |
Figures reflect vendor product briefs, launch materials, and multiply-corroborated industry reporting current as of this writing. Where a vendor has not disclosed a figure or its precision basis in the sources reviewed, that is stated rather than estimated.

What the SKU Variance Actually Means
Qualcomm’s own flagship figure nearly doubled in eighteen months — 45 to 80–85 TOPS (depending on SKU) between Snapdragon X Elite and X2 Elite. Intel’s “48 TOPS” is a flagship-SKU number; the same Core Ultra 200V family ships parts rated at 40 and 47 TOPS, so the marketing name on the box tells a buyer less than the ARK part number does. AMD’s “Ryzen AI 400” spans a 50 TOPS desktop chip and a 60 TOPS mobile flagship under one brand. Apple isn’t chasing any of this — the Copilot+ threshold is a Microsoft ecosystem requirement, not a cross-platform benchmark — and Apple’s M5 announcement doesn’t publish a Neural Engine TOPS figure at all, shifting the AI marketing to GPU-based “Neural Accelerators” instead. A lower or absent TOPS number on a Mac isn’t a deficiency against a spec Apple was never trying to meet.
For a single laptop buyer, that SKU variance inside one marketing family is a minor annoyance — check the exact model number before assuming the flagship spec. For an OEM speccing a laptop line, or a systems integrator quoting a fleet refresh, it’s closer to a bill-of-materials problem. “Core Ultra 200V” and “Ryzen AI 400” aren’t single part numbers you can lock into a purchase order and expect uniform behavior across every unit shipped under that name. Two machines built to the same marketing family can carry different NPU tiers, different Copilot+ headroom, different thermal budgets — the kind of variance a hardware buyer normally expects called out explicitly on a datasheet, not folded into a family name. Anyone qualifying a design against Copilot+ certification, or promising a fleet identical AI behavior, needs the exact part number. The marketing tier isn’t enough to build a purchasing decision on.
Where AI PC Marketing Gets Ahead of the Silicon
Three practices distort the comparison a buyer is actually making, and all three are documented rather than merely suspected.
Combined Platform TOPS vs. NPU-Only TOPS
Microsoft’s 40 TOPS requirement applies to the NPU specifically — it is not a system-wide sum. But vendors routinely promote total platform AI performance, CPU plus GPU plus NPU together, as a single headline figure, and a combined figure like that is a sum of theoretical accelerator peaks — it doesn’t establish that a real application can sustain that aggregate throughput simultaneously, since the three blocks share a single memory bus. Intel’s Lunar Lake launched with a stated 120 TOPS total platform figure: 48 from the NPU, 67 from the Xe2 GPU, and roughly 5 from the CPU — that last figure reported by a single source and less firmly established than the 120 TOPS total, which is corroborated across six independent outlets. Qualcomm has stated a comparable full-platform total near 75 TOPS for the original Snapdragon X Elite, which the company describes as spanning the CPU, GPU, NPU, and a separate always-on micro-NPU block — a fourth compute element the industry’s “combined TOPS” framing typically glosses over entirely.
Intel’s Panther Lake pushed the combined-figure practice further at CES 2026, quoting up to 180 total platform TOPS. Intel’s own materials, corroborated by multiple independent technical outlets, itemize this as 10 from the CPU, 50 from the NPU, and 120 from the GPU — a breakdown that sums correctly. All of these are real numbers describing real silicon. None of them is the number Copilot+ certification is gated on, and a buyer comparing “180 TOPS” against “50 TOPS” without knowing which is which is comparing two different measurements as if they were one.
Precision Basis Isn’t Always Disclosed
TOPS figures are typically quoted at a specific numerical precision — commonly INT8, sometimes INT4 — because lower precision packs more operations per cycle onto the same silicon. Intel has been explicit that its Lunar Lake NPU is rated at INT8 while also supporting higher-precision FP16, and Qualcomm’s own materials confirm the Snapdragon X2 Elite’s 80 TOPS figure is likewise an INT8 rating. Intel has also claimed, via press coverage of its own briefings, that AMD’s and Qualcomm’s contemporaneous NPUs top out at INT8 — a claim about rival silicon relayed through reporting rather than confirmed against AMD’s or Qualcomm’s own documentation, worth reading as Intel’s characterization rather than settled fact. What’s clear regardless: most vendors reviewed here don’t state the precision basis behind their headline figure as plainly as Intel and Qualcomm do for these two chips, which means two “48 TOPS” or “50 TOPS” ratings from different vendors aren’t guaranteed to be measuring the same thing.
Peak Rating vs. What the Silicon Actually Delivers
This is the one that matters most, and it now has a concrete, independently corroborated example behind it. Apple’s M4 Neural Engine is rated at 38 TOPS. Apple’s own materials describe this only as “38 trillion operations per second” and never state a precision basis; the INT8 label commonly attached to it comes from industry press and independent researchers, not Apple. Independent hardware research that bypassed Apple’s CoreML framework to dispatch directly to the Neural Engine found that INT8 weights are dequantized to FP16 before the multiply runs on general matrix workloads — the engine doesn’t execute general compute natively at the INT8 rate its rating implies. Measured throughput on that basis: roughly 19 TFLOPS against a 38 TOPS INT8 headline. That comparison rests on a specific industry convention — counting an INT8 TOPS figure at twice the equivalent FP16 TFLOPS rate — not a direct unit match between two different kinds of arithmetic; under that convention, the measured figure lands at almost exactly half the rated one. A second paper studying the same hardware cites and confirms the figure; a third, broader hardware characterization documents the identical dequantize-before-multiply behavior across several generations of Apple silicon — A14, M1, M2 — with its own measured latency ratios. A separate, more recent measurement from the same research project adds nuance rather than contradicting this: for specific bandwidth-heavy convolution workloads, using INT8 activations to halve the data moved between internal SRAM tiles produced a measured 1.85–1.88x end-to-end speedup — real, but explicitly attributed by the source to reduced memory traffic, not to the compute array running INT8 math natively faster. A separate study of actual on-device LLM decoding on the same hardware found INT8-hybrid decode at roughly parity with FP16 — no speedup at all for that workload — underscoring that the benefit is highly workload-dependent, not a blanket doubling. The precision mode still saves real memory bandwidth, since smaller weights move faster, and that part of the claim holds. What doesn’t hold is the assumption a TOPS rating invites: that INT8 buys a proportional compute-speed multiplier. On this hardware, measured directly, it doesn’t. This same TOPS-measurement problem shows up well beyond laptops — for the fuller picture across phones, cars, and industrial edge silicon, see Edge AI Chips: The Future of AI Hardware and Why They’re Replacing Cloud-Based Intelligence.

Getting the Model Onto the NPU in the First Place
There’s a bottleneck upstream of both TOPS and memory bandwidth that gets even less coverage than either: whether the software stack actually routes a given operation to the NPU at all. A recent measurement study of Apple’s Neural Engine found that placement is a property of how a computation is expressed, not what it computes — a fused RMSNorm operation was fully NPU-eligible, while an arithmetically identical version of the same computation, decomposed into separate steps, ran on the CPU instead. The model didn’t change; the way it was written did, and that alone decided which silicon handled it. The equivalent concern on Windows laptops runs through Windows ML, ONNX Runtime, hardware-vendor execution providers, and how effectively a vendor’s compiler and runtime fuse, quantize, and map a model graph onto the NPU — this extension to the Windows ecosystem is this article’s own inference from the Apple-specific finding, not something the cited research measured directly. Either way, a chip’s rated TOPS describes what its compute array could theoretically do; it says nothing about how much of any specific model’s workload the runtime actually succeeds in routing there instead of falling back to the CPU.
The Real Constraint: Memory Bandwidth, Not Compute
TOPS is a compute-array metric, and inference performance isn’t governed by compute alone. Depending on model architecture, batch size, sequence length, and how well an operation maps to the hardware’s native execution path, a workload can be compute-bound, memory-bandwidth-bound, on-chip-cache-bound, or bound by how the software stack dispatches it — the standard roofline-model picture of accelerator performance, and no single one of those constraints is universally “the” bottleneck across every workload. For the specific class of workloads this piece is about — sustained, often single-sequence, low-batch inference running continuously in the background on a laptop — memory movement is frequently the dominant limiter well before peak arithmetic throughput is reached, because a compute engine can sit idle waiting on data even while its theoretical peak looks excellent on paper. This is ordinary interconnect engineering: shorter, wider, better-controlled paths support higher data rates within signal-integrity margin; longer paths and shared, contended buses cap what a compute block can sustain, independent of how fast that block could run in isolation.

Intel’s own product brief for the Core Ultra Mobile Processors (Series 2) confirms the memory-bandwidth claim directly and in its own words: under the NPU’s AI-compute specifications, Intel states “Up to 2x Bandwidth compared to previous generation” alongside the 48 TOPS figure — a primary-source confirmation, not press paraphrase. The more specific “doubled DMA capacity” wording, describing the NPU’s internal DMA engine rather than bandwidth generally, still traces only to Tom’s Hardware’s detailed account of Intel’s technical briefing, not to this consumer-facing brief. Independent technical testing complicates that DMA-specific claim: one respected independent hardware-analysis outlet measured Lunar Lake’s DMA engines directly and found them weaker than the prior generation’s and weaker than AMD’s competing Strix Point, coming nowhere near saturating available memory bandwidth — a genuine tension this piece can’t resolve, even though the broader bandwidth-doubling claim is now primary-confirmed. Either way, that’s a chipmaker disclosing that a bigger compute array wasn’t sufficient by itself for the workloads it was targeting — the path feeding it needed widening too. Whether Intel’s separate, system-level decision to move memory onto the package was made for the same reason is this article’s inference, not a claim Intel has made explicitly anywhere reviewed in this research; Intel’s own public framing for on-package memory centers on power draw and motherboard footprint.
Thermal Budget Compounds the Problem
A peak TOPS figure is measured under best-case conditions — there’s no standardized, audited cross-vendor benchmark methodology enforcing a common precision, duration, or concurrent-load standard, so not every vendor’s published figure was necessarily generated the same way. What’s consistent across vendors is that peak figures are measured without the CPU and GPU simultaneously drawing on the same power and memory bus a real multitasking session does, and a thin chassis has one thermal ceiling shared across all three blocks. Independent research on sustained on-device inference gives the general mechanism a measured shape, though not on an NPU: a preprint found an iPhone 16 Pro losing nearly half its throughput within two inferences on a sustained language-model decode task, settling about 44 percent below its peak — but that workload ran on the phone’s GPU via Apple’s MLX framework, which the paper’s own methodology states explicitly does not target the Neural Engine. So this is evidence that a mobile accelerator can throttle hard under sustained AI load, not evidence about NPU throttling specifically. No equivalent measured dataset — on a GPU or an NPU, phone or laptop — exists for any of the Intel, AMD, Qualcomm, or Apple silicon this article actually covers. The mechanism is well-established engineering fact independent of this specific study; the size of the gap on any of this article’s specific chips remains unmeasured in anything public.

None of this makes the peak TOPS figure fake. It makes it a ceiling for a specific, workload-dependent bottleneck — one Intel’s own engineering choices suggest the chipmakers take seriously, and one a real laptop, under real thermal load, may not hold for long on the workloads where memory movement dominates.
What to Actually Evaluate When Buying an AI PC
None of this is a reason to write off the category. The 40 TOPS gate is a legitimate floor for a defined set of features, and every current Windows Copilot+ processor family discussed here includes SKUs that meet or exceed the requirement. The mistake is treating the number as a forecast of how the machine will feel, rather than a minimum bar for eligibility.
What actually determines the experience:
- Which specific on-device feature you’ll use — live translation, Recall-style semantic search, background effects — and whether that feature’s documented requirement is met, not exceeded by some multiple.
- Whether a comparison cites NPU-only TOPS or a combined platform figure, and at what precision, before you set it against another vendor’s number.
- Which exact SKU a laptop ships, not just the marketing family name — “Core Ultra 200V” and “Ryzen AI 400” each span a real range of NPU ratings under one badge.
It’s also worth flagging a limit on how far this concern actually extends: Microsoft’s own Copilot+ hardware certification qualifies specific OEM devices before they can carry the badge, and whether that process validates any sustained-performance threshold — not just a peak spec — wasn’t confirmed in this research. If it does, a meaningful share of the peak-versus-sustained gap described above may already be addressed in certified retail systems specifically, even though the gap in the underlying silicon remains real.
The engineering underneath the “AI PC” label is real. As NPU compute scales, the memory subsystem, thermal budget, and software stack around it increasingly determine how much of that theoretical throughput reaches a real workload — that gap, not the TOPS number, is what will decide whether a given “AI PC” delivers on what its spec sheet promises.
This analysis draws on twenty-plus years in PCB manufacturing and technical sales, where a spec sheet is only ever the beginning of a conversation — every rated figure in electronics eventually meets a bill of materials, a yield curve, and a thermal budget that decide whether it holds up in a real product. The same discipline applies here: a TOPS rating, like a component’s datasheet current rating, is a starting point for verification, not a conclusion.
About the Author
Imran Valiani | Sales Director, PCB Electronics Manufacturing
20+ years working with major Bay Area and global tech clients. Founder of Silicon to Software, where I write about the hardware layer — PCB fab, AI gear, autonomous systems, and cyber — the stuff most tech writers have never touched. Literally.
Follow: X @SiToSoftware | LinkedIn
This article was developed with AI assistance and edited, fact-checked, and reviewed by the author. See my full AI disclosure.
Sources
Company and standards documentation
- Microsoft, “Windows 11 Specs and System Requirements” and Microsoft Learn, “Copilot+ PCs developer guide” — Copilot+ hardware requirements
- Microsoft, “Copilot+ PCs & Windows PCs: Differences?” — feature gating and on-device AI framing
- Intel Corporation, “Product Brief: Intel Core Ultra Mobile Processors (Series 2)” — Lunar Lake NPU TOPS, bandwidth, on-package memory
- Intel Corporation, Panther Lake / Core Ultra Series 3 launch materials (CES 2026, intc.com investor-relations press releases)
- Qualcomm, Snapdragon X Elite, X Plus, and X2 Elite product briefs (qualcomm.com)
- AMD, Ryzen AI 300 and Ryzen AI 400 series newsroom announcements and product pages (amd.com)
- Apple Newsroom, “Apple introduces M4 chip” (May 2024) and “Apple unleashes M5” (October 2025)
Independent research
- “Inside the M4 Apple Neural Engine” series, including benchmark data at github.com/maderix/ANE — original reverse-engineered ANE measurements
- “Orion: Characterizing and Programming Apple’s Neural Engine for LLM Training and Inference,” arXiv 2603.06728
- “Apple Neural Engine: Architecture, Programming, and Performance,” arXiv 2606.22283
- “What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Engine,” arXiv 2608.22110
- “LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load,” arXiv 2603.23640
Industry reporting (used for corroboration where primary documentation was unavailable or incomplete)
- The Register, Tom’s Hardware, Engadget, Forbes/TIRIAS Research — Intel Lunar Lake and Panther Lake architecture coverage
- Notebookcheck — AMD Ryzen AI 300/400 and Qualcomm Snapdragon X2 Elite SKU-level reporting
- chipsandcheese.com — independent technical testing of Intel Lunar Lake NPU DMA performance
Note: not every claim in this article carries equal evidentiary weight — several points are explicitly flagged in the text as company-stated claims, single-study findings, or this publication’s own engineering inference rather than independently verified fact. Where that distinction matters, it’s called out inline.