Beginners think machine learning is about models; people who have shipped it know it is mostly about data and the pipeline that moves it. The training call — model.fit(...) — is a couple of lines and a small fraction of the work. Everything around it, the collecting and cleaning and splitting and evaluating and monitoring, is where projects are actually won or lost. It is also where the subtle, self-inflicted failures live: pipelines that look like they are succeeding while quietly measuring the wrong thing. This is a tour of the whole pipeline, stage by stage, with the traps marked.
The pipeline at a glance
Notice that it is a loop, not a line: monitoring feeds back into new data and retraining. Let us walk it.
Ingestion and data quality
Everything starts with getting data in — from databases, logs, files, sensors, third-party feeds — and it arrives messier than you expect: missing values, duplicates, inconsistent units, wrong types, encoding quirks. The unglamorous work of cleaning and validation is foundational, because a model faithfully learns whatever is in its data, including the mistakes. "Garbage in, garbage out" is not a slogan here; it is the single most common reason a model underperforms. Validating incoming data against expectations — ranges, types, allowed values — catches problems before they poison training.
Feature engineering
Feature engineering turns raw data into the inputs a model can actually use: encoding categories as numbers, scaling values to comparable ranges, deriving informative combinations, extracting signal from dates or text. This stage is frequently where the largest gains hide. A better feature — one that exposes the real structure of the problem — often beats a fancier model on the same raw data. It rewards domain knowledge as much as ML knowledge, because knowing which derived quantity matters is a question about the problem, not the algorithm.
The split — and the cardinal rule
Before training you divide data into a training set (to learn from), a validation set (to tune choices), and a test set (an honest final estimate of performance on unseen data). And here is the rule that beginners violate most often, with expensive consequences:
Split first, then preprocess — fit only on training data
Any transformation that learns from the data — scaling, imputing missing values, selecting features — must be fit on the training set alone, then applied to validation and test. If you fit it on the full dataset before splitting, information from the test set leaks into training, and your evaluation becomes optimistic and wrong. This is called data leakage, and it is worth its own article; the discipline is to split first, always.
The clean way to enforce this is a pipeline object that bundles preprocessing and the model together, so that when you cross-validate, every fold re-fits the preprocessing on just that fold's training portion. Tooling like scikit-learn's Pipeline exists precisely to make the correct thing the easy thing.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
# Scaler is fit ONLY on the training fold each time — no leakage from test.
pipe = Pipeline([
("scale", StandardScaler()),
("model", LogisticRegression()),
])
pipe.fit(X_train, y_train)Training and evaluation
Training fits the model to the training data; evaluation asks whether it learned anything that generalises. The trap here is trusting a single number. Accuracy alone can be actively misleading — on an imbalanced problem where 99 percent of cases are one class, a model that always predicts that class is 99 percent accurate and completely useless. Which metric to trust depends on what mistakes cost: precision and recall when false positives and false negatives differ in cost, and always an evaluation on data the model has never seen.
Deploy and monitor
Shipping the model is not the finish line. The world the model was trained on keeps changing — user behaviour shifts, inputs drift, an upstream data source changes format — and a model that was accurate at launch can silently decay. This is why monitoring belongs in the pipeline: track input distributions and output quality in production, alert when they drift, and retrain on fresh data when they do. There is a well-known observation in the field that the machine-learning code is a small box in a much larger system of data collection, feature extraction, serving, and monitoring — and that the surrounding infrastructure is where long-term maintenance cost, sometimes called technical debt, accumulates.
Practical takeaway
If you are new to building ML systems, redistribute your attention: spend less of it choosing architectures and more of it on the pipeline. Get the data clean and validated, engineer features that expose the real signal, split before you preprocess and bundle preprocessing into a pipeline so leakage cannot sneak in, evaluate on unseen data with a metric that reflects real costs, and plan from day one to monitor and retrain. The model is the part everyone talks about; the pipeline is the part that determines whether the thing actually works six months after launch.
Sources & Further Reading
- 01Rules of Machine Learning: Best Practices for ML Engineering — Martin Zinkevich (Google)Hard-won practical guidance on building real ML systems.
- 02Common pitfalls and recommended practices — scikit-learnIncluding how preprocessing before splitting leaks data.
- 03Pipelines and composite estimators — scikit-learnHow to bundle preprocessing with a model to prevent leakage during cross-validation.
Editorial note — A conceptual explainer of the standard ML data-pipeline stages and their common failure modes. Code is illustrative; no benchmark numbers are quoted, and the imbalanced-accuracy example (99 percent) is a generic illustration, not a measured result.

