Research
Systems for Machine Learning
The infrastructure machine learning runs on — data pipelines, training and serving efficiency, and the operational side of models in production.
A model is a small part of a machine learning system. The rest — ingestion, labelling, feature computation, retraining, monitoring, serving under a latency budget — is where most of the cost and most of the failures live.
We study that layer: how to build data pipelines whose quality can be measured, how to serve models efficiently, and how to keep a deployed system honest as its inputs change. Much of this comes from practice; the lab’s work in this area is informed by industrial MLOps deployments.