A neural processing unit (NPU) is a processor designed to run the math behind AI models, mostly the matrix calculations used in neural networks, quickly and with little power. NPUs are built into many newer laptop and desktop processors (the basis of the AI PC category), smartphones, cameras and other edge devices. They are mainly used to run trained models, not to train them.
At a glance
- An NPU is a specialized chip, or part of a chip, designed to run AI workloads efficiently at low power.
- It is mainly used for AI inference: running trained models for tasks like speech recognition, video effects, image analysis and small language models.
- It usually sits alongside a CPU and GPU, and software decides which one handles each task.
- Performance is often quoted in TOPS (trillions of operations per second), which vendors measure in different ways.
- Its value depends on software support: an NPU that no application uses adds little.
What problem it solves
AI models do huge numbers of similar calculations. A general-purpose CPU can run them, but slowly and with high power draw. A GPU can run them quickly but uses more power than a laptop or camera may want to spend on a background task. An NPU is designed for this one kind of work, so it can run common AI features continuously, such as noise suppression in a meeting, without draining the battery or slowing other applications.
For buyers, NPUs matter because they shift some AI work from the cloud to the device. That can reduce per-use cloud costs, keep working when the connection drops and keep some data local. It also affects which devices can run features that operating systems and software vendors are adding.
How it works
An NPU contains many small processing units arranged to run the multiply-and-add operations at the heart of neural networks in parallel, often using lower-precision numbers that are good enough for running a trained model. It is typically built into the same processor package as the CPU and GPU in a PC or phone, or into a system-on-chip in a camera or industrial device.
Software reaches the NPU through drivers and AI runtimes provided by the chipmaker or operating system. Developers, or the operating system itself, decide which tasks go to the NPU. Models usually need to be prepared or converted for a particular NPU, which is why software support differs from one chip family to another.
Chipmakers use different names for their NPUs, and headline TOPS figures are not directly comparable across vendors. Real results depend on memory bandwidth, model size, precision and software.
When it matters for buyers
- During a PC refresh. If operating system or software features you want require an NPU, it becomes part of the device specification. Plan this into your hardware refresh cycle.
- For edge and IoT projects. Cameras, gateways and devices with NPUs can analyze video or sensor data locally, which ties into edge computing designs and can reduce how much data you send to the cloud. See our edge compute overview for where that fits.
- When comparing spec sheets. Treat TOPS as a rough guide and ask for results on the workloads you actually run.
- When planning heavy AI work. NPUs in endpoints are not a substitute for data center GPUs or GPU as a service (GPUaaS) for training or serving large models.
Questions to ask vendors
- Which NPU is in this device, and what software and AI runtimes support it today?
- Which of our applications will use the NPU, and which tasks still run on the CPU, GPU or in the cloud?
- How was the TOPS figure measured, and can you show results for workloads like ours?
- What memory does the device need to run on-device models well?
- How are NPU drivers and AI runtimes updated, and for how long?
- For edge devices: which models can run locally, and how are they deployed and updated across the fleet?
How it differs from a GPU
A graphics processing unit (GPU) is a general parallel processor. It handles graphics and is the main workhorse for training and serving large AI models in data centers, which is why cloud providers sell it as GPU as a service (GPUaaS). An NPU is more specialized: it is designed mainly to run trained models efficiently at low power, often inside a laptop, phone or edge device. A GPU can typically handle bigger and more varied AI jobs; an NPU is typically better suited to small, sustained, always-on tasks where battery life and heat matter. Many devices include both and split work between them.
