A small language model (SLM) is an AI model that understands and generates text like a large language model (LLM), but with far fewer parameters, the internal values a model learns during training. That smaller size means it can run on a single server, a workstation or even a laptop or phone, respond faster and cost much less per request. The trade-off is narrower knowledge and weaker performance on complex, open-ended tasks, so SLMs are usually aimed at focused jobs such as classifying tickets, extracting fields from documents or answering questions about a defined set of content.
At a glance
- “Small” is relative: an SLM is much smaller than the largest models of its time, with no fixed size threshold.
- Smaller models usually cost less to run and respond faster, and many can run on your own hardware or on devices.
- They tend to do well on narrow, repeatable tasks and less well on broad reasoning or general knowledge.
- Many are fine-tuned on an organization’s own examples to improve accuracy on a specific job.
- Licences, accuracy and security still need checking, just as with larger models.
What problem it solves
Many organizations start with AI through large, general-purpose models delivered as a cloud service and billed per token, the small chunks of text a model reads and writes. That works well for drafting and open-ended questions, but it can become expensive at volume, add latency that hurts real-time use, and require sending data to a third party.
A lot of business AI work is narrow: sort this email, pull these fields from an invoice, tag this support ticket, summarize this call note in a set format. A smaller model can often do that kind of task well enough at a fraction of the compute. Because it needs less hardware, it can also run inside your own environment, at a branch site or on a user’s device, which helps when data must stay on premises or connectivity is unreliable.
How it works
Training and distillation. SLMs are trained the same basic way as larger language models, on large amounts of text, but with a smaller architecture. Many are produced by distillation, where a smaller model learns to imitate the outputs of a larger one, or by pruning and compressing a larger model.
Specialization. A general SLM is often fine-tuned on examples of the specific task it will do, such as labelled support tickets. This is where much of its value comes from: a small model trained for one job can approach the accuracy of a much larger general model on that job.
Deployment. SLMs run through the same AI inference process as other models. Depending on size, they may run on a modest GPU server, a CPU, a workstation, or a phone or PC with an AI accelerator (often called a neural processing unit, or NPU). Quantization, storing the model’s numbers at lower precision, can shrink memory use further at some cost to accuracy.
Grounding and routing. SLMs are often paired with retrieval-augmented generation (RAG) so they answer from your documents instead of their limited built-in knowledge. Some systems route simple requests to a small model and hand harder ones to a larger model, balancing cost and quality.
When it matters for buyers
- When AI costs climb. If usage-based AI bills grow with volume, moving well-defined, high-volume tasks to a smaller model can cut cost per request.
- When data has to stay inside. Running a model on your own infrastructure is one route to private AI, useful for regulated or confidential data.
- When speed matters. Real-time uses such as live call assistance or in-app suggestions benefit from faster responses.
- When working at the edge. Sites, vehicles or devices with limited connectivity may need a model that runs locally; see our edge compute options.
- When a vendor says “we use our own model.” Many AI features in business software run on smaller, task-specific models; ask what that means for accuracy and data handling.
Questions to ask vendors
- Which model does the feature use, how large is it, and where does it run?
- How does its accuracy on our task compare with a larger model, and can we test it on our own data?
- Was it fine-tuned, and on what data? Is any of our data used for further training?
- What hardware does it need if we host it ourselves, and what are the licence terms for commercial use?
- How do you handle requests the small model can’t answer well: does it escalate to a larger model or a person?
- How is it priced: per request, per token, per device or per user?
- How will the model be updated, and how will we be told when it changes?
How it differs from a large language model (LLM)
An LLM is built for breadth: wide general knowledge, flexible reasoning and the ability to handle many kinds of request, at the cost of heavy compute, usually delivered from a cloud data center. An SLM gives up some of that breadth for lower cost, faster responses and the option to run locally. Both are language models and both can make mistakes; the choice depends on the task. Many organizations end up using both: a large model for open-ended work and smaller models for high-volume, well-defined steps. Some large and small models are versions of the same foundation model family, released at several sizes.
