Machine learning operations (MLOps) is the set of practices, roles and tools that take machine learning (ML) models from experiment to dependable production use, and keep them working afterward. It covers how data is prepared and versioned, how models are trained, tested and approved, how they are deployed, and how their accuracy, cost and behaviour are monitored over time. MLOps applies ideas from DevOps, such as automation and repeatable releases, to the extra challenges of software that learns from data.
At a glance
- MLOps covers the full model lifecycle: data, training, testing, deployment, monitoring and retraining.
- It exists because ML models can degrade as real-world data changes, unlike ordinary code that behaves the same until someone changes it.
- Core practices include versioning data and models, automated pipelines, model registries and production monitoring.
- It supports governance and audit by recording which model, trained on which data, produced a given result.
- It matters most when you build, fine-tune or host models yourself; with SaaS AI features, the vendor carries most of it.
What problem it solves
Many organizations can build a promising model in a notebook, but far fewer get it into production and keep it working. Common causes are manual hand-offs between data scientists and IT, no repeatable way to retrain, unclear ownership once the model is live, and no way to tell when accuracy has slipped.
Machine learning brings problems that ordinary software doesn’t. A model’s behaviour depends on its training data, so the data needs to be tracked as carefully as code. Accuracy can drift as customers, products or fraud patterns change. Regulators, auditors and customers may ask why a model made a decision, which requires knowing exactly which version was running. MLOps gives teams a repeatable process for all of this, so models can be updated safely and their behaviour can be explained after the fact.
How it works
Data management. Training data is collected, cleaned, labelled and versioned, often from a data lake or warehouse, with data governance rules on quality, privacy and access.
Experiment tracking. Each training run records the data, code, settings and results, so teams can compare approaches and reproduce a model later.
Model registry and approval. Trained models are stored in a registry with their version, documentation, test results and approval status. Higher-risk models may need sign-off before release.
Automated pipelines. Like continuous integration and continuous delivery (CI/CD) for software, pipelines retrain, test and deploy models in a repeatable way, reducing manual errors.
Deployment and serving. Models are released for AI inference as APIs, batch jobs or embedded components, often with staged rollouts to compare a new version against the current one.
Monitoring and retraining. In production, teams watch accuracy, data drift, bias, latency and cost, and trigger retraining or rollback when measures cross agreed thresholds.
For generative AI and foundation models, similar practices are sometimes called LLMOps, adding prompt versioning, output evaluation and token cost tracking.
When it matters for buyers
- When AI pilots stall. If models work in testing but never reach production, the missing piece is often MLOps process and tooling, not a better model.
- When choosing an ML platform. Cloud AI platforms and specialist tools bundle MLOps features differently; compare on your team’s actual workflow.
- When regulations or audits apply. Model versioning and decision logs support AI governance and requests to explain automated decisions.
- When outsourcing model work. If a partner builds models for you, agree who runs monitoring and retraining after go-live.
- When budgeting. Training and serving costs, especially on GPUs, need tracking just like other cloud spend.
For help sourcing data and AI platforms, see our analytics and business intelligence options.
Questions to ask vendors
- How does your platform version data, code and models together so we can reproduce a result?
- What monitoring is included for accuracy, drift, bias, latency and cost, and how are we alerted?
- How are model approvals and rollbacks handled, and are they logged for audit?
- Which deployment targets do you support: your cloud, our cloud account, on-premises or edge?
- How do you support generative AI models, such as prompt versioning and output evaluation?
- If your team builds the model, who monitors and retrains it after launch, and at what cost?
- Can we export our models, metadata and pipelines if we change platforms?
How it differs from DevOps
DevOps is a set of practices for building, testing and releasing software quickly and reliably by bringing development and operations teams together. MLOps applies the same ideas to machine learning but deals with extra moving parts: training data, experiments, model versions and accuracy that changes over time without any code changes. Many organizations run MLOps alongside their DevOps and CI/CD tooling, sharing infrastructure while adding ML-specific steps such as data validation, model evaluation and drift monitoring.
