Designs by Duhart All writing

·6 min read·machinelearning · datascience · python · scikitlearn · ai · mlengineering · codinginterview · softwareengineering · statistics · learntocode

5 ML Algorithms You Should Already Know (With scikit-learn Code)

Linear regression, logistic regression, random forests, k-means and gradient boosting: the intuition, when to reach for each, the gotcha that bites, and a scikit-learn snippet for every one.

Cover slide: 5 ML Algorithms You Should Already Know (With scikit-learn Code). Bottom left, an "Interviewing at" badge with the logos of Google, Meta, Amazon, Spotify.

Why the classics still matter

I posted a short version of this on Instagram a while ago. This one has code for every algorithm, the gotcha I'd ask about in an interview, and a runnable script at the end.

Here's the thing people skip when they jump straight to deep learning: most of the data a business owns is a table. Customers and orders, sessions and clicks, sensors and readings. On tables, five classical algorithms do most of the real work, they train in seconds, and you can explain them to the person who has to sign off on the result.

They're also the best way to find out whether you need anything bigger. A ten line baseline gives you a number. If the fancy model can't beat it by a margin that matters, you just saved yourself a GPU bill.

A personal example: the ranking model I'm designing for my own personalization service is LightGBM, which is gradient boosted trees, one model per region and product, scored inside a Go service. Not a neural network. The inputs are tables, so the classics win again.

All the snippets use scikit-learn and assume you already have X_train, X_test, y_train and y_test. The full script at the end builds those from a dataset that ships with scikit-learn, so you can run everything.

5 ML algorithms you should already know

Cover slide: 5 ML Algorithms You Should Already Know (With scikit-learn Code). Bottom left, an "Interviewing at" badge with the logos of Google, Meta, Amazon, Spotify.
The carousel this article expands on.

5 ML algorithms in 40 seconds

The video version: why the classics still win, and all five in one line each.
Opening frame of the explainer video. A 39 second explainer in the house dark style with burned-in captions: a joke about training a neural network on 5,000 rows, why tables favour classic algorithms, the five algorithms in one line each, a four line random forest baseline, why 99 percent accuracy can be useless, and an end card reading Save this, designsbyduhart.org.
The video version: why the classics still win, and all five in one line each. Watch the video: https://designsbyduhart.org/blog/ml-algos-5/

Why these five still win

Slide 2 of 10: WHY THESE FIVE STILL WIN. Most real business data is a table. On tables, these five do most of the work. They train in seconds on a laptop. You can explain them to the person who signs off. A ten line baseline tells you if a bigger model is worth it. Interviewers ask about them by name. From my own work: The ranking model I am designing for my own personalization service is gradient boosted trees, not a neural network. Tables, again.
Most real data is a table, and these five run most tables.

1. Linear regression

Intuition. Draw the straight line (or, with many features, the flat plane) through your data that keeps the squared distance between the line and the points as small as possible. Each feature gets a coefficient: how much the prediction moves when that feature goes up by one.

When to use it. Predicting a number: a price, a delivery time, next week's demand. And as the baseline for any regression problem, because it's fast, it's explainable, and it's surprisingly hard to beat on small, clean data.

The gotcha. Squared error punishes big misses hard, so a single outlier can pull the whole line toward it. And when two features carry the same information (square feet and number of rooms), the coefficients can swing wildly between them even though the predictions barely change. Ridge regression adds a penalty on large coefficients and calms that down.

Linear regression

Slide 3 of 10: LINEAR REGRESSION. Fit the straight line (or plane) that keeps the squared errors smallest. Predicts a number. Why it matters: Use it for prices, demand, durations, and as the baseline every fancier model has to beat. The coefficients tell you how much each feature moves the answer. A code card (python) shows: from sklearn.linear_model import LinearRegression  reg = LinearRegression().fit(X_train, y_train) print(reg.coef_, reg.intercept_) print(reg.score(X_test, y_test))  # R squared Interview trap: Squared error means one outlier can drag the whole line. Correlated features make coefficients swing; Ridge calms them down.
The baseline for numbers. Watch the outliers.

2. Logistic regression

Intuition. Compute a linear score exactly like linear regression, then squeeze it through a sigmoid so it lands between 0 and 1. That number is a probability. Despite the name, it's a classifier.

When to use it. Any yes or no question: will this customer churn, is this transaction fraud, will this user click. It's the first model I'd try, because it's fast, its coefficients are readable, and you get a probability you can threshold however the business needs.

The gotcha. Two of them. First, scikit-learn's LogisticRegression is regularised by default, which means the scale of your features changes the answer. Put a StandardScaler in front of it. Second, class imbalance. If 99 percent of your rows are "not fraud", a model that always says "not fraud" is 99 percent accurate and useless. Look at precision and recall or ROC AUC, consider class_weight="balanced", and pick the decision threshold on purpose. Nothing says it has to be 0.5.

Logistic regression

Slide 4 of 10: LOGISTIC REGRESSION. A linear score squeezed through a sigmoid into a probability. Despite the name, it classifies. Why it matters: The first thing to try for yes or no questions: churn, fraud, click. Fast, explainable, and it gives you a probability, not just a label. A code card (python) shows: clf = make_pipeline( StandardScaler(), LogisticRegression(class_weight="balanced")) clf.fit(X_train, y_train) p = clf.predict_proba(X_test)[:, 1] Interview trap: With 99% negatives, a model that always says no is 99% accurate. Use precision, recall or ROC AUC, and pick the threshold. 0.5 is not sacred.
A probability for yes or no. Accuracy lies on imbalanced data.

3. Decision trees and random forests

Intuition. A decision tree asks a series of yes or no questions about the features ("is income above 40k?") until it reaches an answer. One deep tree can memorise the training data perfectly and generalise badly. A random forest trains hundreds of trees, each on a random sample of the rows and a random subset of the features at every split, and lets them vote. Their individual mistakes are different, so averaging cancels most of them out.

When to use it. When you want a strong model on tabular data with almost no tuning. Forests handle non-linear relationships and interactions between features, they don't need scaled inputs, and the defaults are usually fine.

The gotcha. The built-in feature_importances_ are computed from how much each feature reduced impurity during training, and they're biased toward features with many unique values, like IDs or exact timestamps. Before you tell anyone "feature X matters most", check with permutation_importance on held-out data.

Random forest

Slide 5 of 10: RANDOM FOREST. Hundreds of decision trees, each on a random sample of rows and features, then a vote. Why it matters: One deep tree memorises the training data. Averaging many different trees cancels most of that out. A strong default with almost no tuning. A code card (python) shows: from sklearn.ensemble import RandomForestClassifier  rf = RandomForestClassifier( n_estimators=300, n_jobs=-1, random_state=0) rf.fit(X_train, y_train) print(rf.score(X_test, y_test)) Interview trap: The built-in featureimportances favour features with many unique values. Check with permutationimportance before you believe them.
Many different trees, one vote. A strong default.

4. K-means

Intuition. You have no labels, just points. Pick a number k. Drop k centres, assign every point to its nearest centre, move each centre to the average of its points, and repeat until nothing moves. What you get is k groups of points that are close to each other.

When to use it. Customer segmentation, grouping similar documents, reducing an image to a handful of colours. Anywhere you want to discover structure that nobody labelled.

The gotcha. The algorithm doesn't choose k. You do, so try several and compare them with silhouette scores (the elbow plot is the other common method, and it's often ambiguous). Distances drive everything, so scale your features first, or the one measured in dollars will swamp the one measured in years. And k-means assumes roughly round clusters of similar size. Long, curved or very uneven groups need something like DBSCAN or a Gaussian mixture.

K-means

Slide 6 of 10: K-MEANS. No labels. Drop k centres, assign every point to the nearest, move each centre to its points' average, repeat. Why it matters: Customer segments, grouping documents, compressing colours. Useful when you want to find structure and nobody has labelled anything. A code card (python) shows: from sklearn.cluster import KMeans  Xs = StandardScaler().fit_transform(X) km = KMeans(n_clusters=4, n_init=10, random_state=0) labels = km.fit_predict(Xs) Interview trap: You pick k, not the algorithm. Check silhouette scores, scale the features first, and expect trouble with long or uneven clusters.
Groups without labels. You choose k.

Watch k-means converge

Lloyd’s algorithm run live on synthetic data: assign, move, repeat. The inertia (total squared distance) only goes down.
Final frame of the animated chart. Animation of k-means on 150 synthetic points in three blobs: three centroids start in poor positions, then alternate between assigning points (recolouring them) and moving to the mean of their points, until they settle on the three groups.
Lloyd’s algorithm run live on synthetic data: assign, move, repeat. The inertia (total squared distance) only goes down. Watch the video: https://designsbyduhart.org/blog/ml-algos-5/

5. Gradient boosting

Intuition. Where a random forest builds trees independently and averages them, boosting builds them in sequence. Each new, small tree is trained on the errors the ensemble still makes (technically, on the gradient of the loss), and gets added with a small weight. Hundreds of weak trees, each correcting the last, add up to a very strong model.

When to use it. When accuracy on tabular data matters most. Gradient boosted trees win a large share of tabular competitions and run a lot of production ranking and risk systems. XGBoost, LightGBM and CatBoost are all this idea with clever engineering. In scikit-learn, HistGradientBoostingClassifier is the fast one and handles missing values natively.

The gotcha. Boosting will keep reducing training error long after it stopped learning anything real. Use early stopping on a validation split, and remember the trade: a lower learning rate needs more trees but usually generalises better.

Gradient boosting

Slide 7 of 10: GRADIENT BOOSTING. Trees built one after another, each one trained to fix the mistakes the ones before it still make. Why it matters: Usually the most accurate option on tabular data. XGBoost, LightGBM and CatBoost are all this idea, tuned for speed. A code card (python) shows: from sklearn.ensemble import ( HistGradientBoostingClassifier)  gb = HistGradientBoostingClassifier( learning_rate=0.1, max_iter=500, early_stopping=True) gb.fit(X_train, y_train) Interview trap: Keep adding trees and it will memorise the noise. Use early stopping on a validation split, and trade a lower learning rate for more trees.
Trees in sequence, each fixing the last. Stop early.

Which one, when

If you remember nothing else, remember the order. Start with the simple model for your task, get an honest score, and move down the table only when the baseline says it's worth it.

Which one, when

Slide 8 of 10: WHICH ONE, WHEN. Start at the top of your row. Move down only when the baseline says so. A table: ALGORITHM, TASK, REACH FOR IT WHEN; LINEAR, number, you need a baseline or to explain it; LOGISTIC, yes or no, you want a probability, fast; FOREST, either, you want strong with no tuning; K-MEANS, groups, there are no labels; BOOSTING, either, accuracy matters most.
Start simple and move down only when the baseline says so.

The trap under all five: leakage

The most common way I've seen a model score 99 percent in a notebook and fall over in production is data leakage. The classic version: fit the scaler (or the imputer, or the feature selection) on the whole dataset, then split into train and test. The test set's statistics have now leaked into training, and your score is optimistic.

The fix is structural. Put every preprocessing step inside a scikit-learn Pipeline, and evaluate the pipeline with cross-validation. Each fold then refits the scaler on its own training data only. The other kind of leakage is a feature that secretly contains the answer, like a "cancellation date" column in a churn model, or rows from the future in a time series. If a score looks too good, go looking for that before you celebrate.

The bug that fakes 99%

Slide 9 of 10: THE BUG THAT FAKES 99%. Scale or impute on the whole dataset before you split, and the test set leaks into training. Why it matters: Put every preprocessing step inside a Pipeline. Cross-validation then refits the scaler on each training fold only, so the score is honest. A code card (python) shows: pipe = make_pipeline(StandardScaler(), LogisticRegression()) scores = cross_val_score(pipe, X, y, cv=5) print(scores.mean(), scores.std()) Interview trap: A score that looks too good usually is. Look for leakage first: a feature that is really the label, or rows from the future.
Preprocess inside a Pipeline and cross-validate the whole thing.

All five on datasets that ship with scikit-learn. Runs in a few seconds on a laptop.

python
from sklearn.datasets import load_breast_cancer, load_diabetes
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LinearRegression, LogisticRegression
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier
from sklearn.cluster import KMeans
from sklearn.metrics import roc_auc_score, silhouette_score

# regression: diabetes progression
Xr, yr = load_diabetes(return_X_y=True)
Xr_tr, Xr_te, yr_tr, yr_te = train_test_split(Xr, yr, random_state=0)
print("linear R^2", LinearRegression().fit(Xr_tr, yr_tr).score(Xr_te, yr_te))

# classification: breast cancer
X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)
models = {
    "logistic": make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)),
    "forest": RandomForestClassifier(n_estimators=300, n_jobs=-1, random_state=0),
    "boosting": HistGradientBoostingClassifier(early_stopping=True, random_state=0),
}
for name, m in models.items():
    m.fit(X_tr, y_tr)
    auc = roc_auc_score(y_te, m.predict_proba(X_te)[:, 1])
    cv = cross_val_score(m, X, y, cv=5, scoring="roc_auc")
    print(f"{name:9s} test AUC {auc:.3f}  cv AUC {cv.mean():.3f} +/- {cv.std():.3f}")

# clustering: same features, no labels
Xs = StandardScaler().fit_transform(X)
for k in (2, 3, 4):
    labels = KMeans(n_clusters=k, n_init=10, random_state=0).fit_predict(Xs)
    print("k-means k =", k, "silhouette", round(silhouette_score(Xs, labels), 3))

How I'd answer in an interview

If someone asks which algorithm you'd use, the best first answer is a question about the data: how many rows, what kind of target, does it need to be explained. Then name the baseline you'd start with, the model you'd move to if it fell short, and how you'd measure it without fooling yourself. That answer shows judgement, which is what they're actually testing.

Save it for your next interview

Closing slide 10 of 10: Save this. Recap: Linear regression: numbers, and your baseline; Logistic regression: yes or no, with a probability; Random forest: strong default, little tuning; K-means: groups without labels, you pick k; Gradient boosting: top accuracy on tables; Accuracy lies on imbalanced classes; Preprocess inside a Pipeline, always. Your turn: Which of the five did you use first, and on what problem? Full write-up with code at designsbyduhart.org.
The recap card.

Next in the interview series: gRPC fundamentals, from .proto contracts to streaming and load balancing.

More: LinkedIn · Instagram. Portfolio and case studies: designsbyduhart.org.

If any of this saved you an afternoon, Buy me a coffee.