MLOps (machine learning operations) is the practice of getting machine learning models into production and keeping them reliable. It covers versioning data and models, tracking experiments, automated testing and deployment, serving predictions through APIs, and monitoring models so you notice when they start to go wrong.

Why do we need MLOps?

A model that works in a notebook often fails in the real world. The data changes, code and data versions get mixed up, nobody can reproduce last month's results, and a model quietly gets worse without anyone noticing. MLOps brings the discipline of software engineering to machine learning so that models are reproducible, testable and observable.

The MLOps pipeline, step by step

  1. Version data and code so every model can be traced to exactly what trained it.
  2. Track experiments: parameters, metrics and artefacts for every training run.
  3. Automate training and testing with a pipeline, including checks on data quality and model accuracy.
  4. Package the model in a container with its dependencies.
  5. Deploy and serve it behind an API, with the ability to roll back.
  6. Monitor latency, errors, cost and changes in the input data (data drift).
  7. Retrain when monitoring shows the model's performance falling.

Common MLOps tools

NeedTypical tools
Code and data versioningGit, DVC
Experiment trackingMLflow, Weights & Biases
PackagingDocker
Orchestration and scalingKubernetes
CI/CDGitHub Actions and similar
Model servingFastAPI, TorchServe, Triton, vLLM for LLMs
MonitoringPrometheus, Grafana

MLOps for large language models

LLM applications add new concerns: serving large models efficiently (quantisation, batching, caching), controlling cost per token, evaluating outputs that have no single right answer, and monitoring for unsafe or incorrect responses.

How to learn MLOps

MLOps sits on top of programming, machine learning and basic cloud skills. In Program Zero it comes in Phase 10 (cloud, MLOps and deployment), where you deploy your own LLM and RAG chatbot with a full CI/CD pipeline. Read what RAG is first if you're new to LLM apps.