AI inference is the step where a trained AI model is put to work: it takes new input, such as a typed question, a document, an image or a sensor reading, and produces an output, such as an answer, a classification, a prediction or generated text. Training is how a model learns; inference is how it is used. Every time someone asks a chatbot a question or a camera flags an object, inference is happening somewhere, whether in a provider’s data center, on your own servers or on the device itself.
At a glance
- Inference is running a finished model on new data; training is building the model in the first place.
- It happens every time the model is used, so its cost and capacity needs grow with usage.
- It can run in a public AI service, on rented or owned servers, or at the edge on local hardware and devices.
- Hardware ranges from ordinary CPUs to GPUs and specialized AI accelerators, depending on the model and speed needed.
- Speed (latency), throughput, cost per request and where data goes are the main buyer trade-offs.
What problem it solves
A trained model is just a large file until something runs it. Inference is the operational side of AI: serving the model to users and applications reliably, quickly and at a sensible cost. For most organizations, this is where AI spending and engineering effort actually land, because they use models others have trained rather than training their own.
It is also where the practical questions arise. A customer-facing chatbot needs answers in seconds, not minutes. A factory inspection system may need results in milliseconds and cannot depend on an internet link. A legal team may not want documents sent to an outside service. How and where inference runs decides whether those requirements can be met.
How it works
The model is loaded. A trained model, from a provider or an open model you download, is loaded into memory on the chosen hardware. Large language models can require a lot of memory, which is why they commonly run on GPUs or AI accelerators.
Requests come in. An application sends input to the model, usually through an API. For a large language model (LLM), the input is text broken into tokens; for other machine learning (ML) models, it might be an image, a set of numbers or a transaction record.
The model computes an output. The model processes the input and returns its result. Serving software manages queues, batching of requests and scaling across machines so many users can be served at once.
Where it runs. Options include a hosted AI service, where the provider runs everything; rented capacity such as GPU as a Service (GPUaaS) or bare metal servers; your own hardware in a data center or as private AI; or edge computing locations and devices close to where data is produced. Each changes cost, control, latency and how much operating work falls to you.
Techniques such as using smaller models, compressing models or caching common answers can reduce hardware needs, often with some trade-off in quality.
When it matters for buyers
- When AI usage is growing. Inference is a recurring cost tied to volume; forecast it before rolling a tool out widely.
- When response time matters. Customer-facing, voice and real-time industrial uses need low latency, which may favor nearby or on-site hardware.
- When data can’t leave your control. Self-hosted or private inference keeps prompts and outputs inside your environment, at the cost of running it.
- When deciding whether to rent or own hardware. Steady, high-volume inference can make dedicated capacity worth comparing against per-request pricing.
Our artificial intelligence overview covers hosted and self-managed options, and our edge compute and bare metal overviews cover infrastructure for running models yourself.
Questions to ask vendors
- How is inference priced (per request, per token, per GPU hour), and what usage would our workload generate?
- What response times can we expect at our volume, and what is committed in the contract?
- Where does inference run, and do prompts, outputs or logs leave that location?
- Which models and model sizes can we use, and can we bring our own?
- What hardware is behind the service, and is capacity shared with other customers or reserved?
- How does the service scale at peak, and what happens when capacity runs short?
How it differs from AI training
Training builds a model: it processes very large datasets over hours, days or longer, often across many GPUs working together in high-performance computing (HPC) clusters, and it happens occasionally. Inference uses the model: it handles individual requests continuously, and its demands depend on how many people and systems use the model. Training is usually a project cost borne by whoever builds the model; inference is an ongoing operating cost borne by whoever uses it. Most mid-market organizations pay for inference, directly or inside a subscription, and rarely for training large models.
