Skip to content

Machine Learning Explained: Supervised, Unsupervised, RL

Tested with: Python 3.12.3, scikit-learn 1.9.1, NumPy 2.5.3, SciPy 1.18.1 (CPU). Last run 2026-09-27.

A machine learning model doesn't get told the rules. It gets shown enough examples that it infers them on its own. That's the entire difference from traditional software, where a programmer writes the logic explicitly. A spam filter isn't running a list of "if email contains X, mark as spam" rules a person wrote, it learned what spam tends to look like from large numbers of labeled examples of spam and not-spam. Real filters still mix in hand-written rules and blocklists. The learned model is the part that keeps up when spammers change tactics. Product recommendations and speech recognition work on the same principle.

Labels, no labels, or rewards

Which of the three you use depends on what your data looks like: labeled, unlabeled, or generated by the model's own actions.

Supervised learning is the most common in production systems: you give the algorithm labeled examples (this email is spam, this one isn't) and it learns to predict the label for new, unseen data. This requires labeled data to exist in the first place, and getting those labels is often the slowest part of a real project, slower than the modeling.

Unsupervised learning skips the labels entirely. The algorithm looks for structure in the data on its own, which is how customer segmentation or anomaly detection happens without anyone pre-tagging what "normal" looks like. The trade-off: without labels, there's no correct answer to check the output against. Deciding whether the clusters it found are useful takes human judgment in a way supervised learning doesn't.

Reinforcement learning works differently. Instead of a fixed dataset, an agent takes actions in an environment and learns from reward signals, adjusting its behavior to maximize cumulative reward over time. DeepMind's AlphaGo, which beat Lee Sedol at Go in 2016, paired it with tree search, and a growing share of robotics research uses it too. It usually needs far more samples and compute than the other two, because the agent has to try things (bad moves included) to find out what works, often in a simulator. The reward can also be delayed: a move early in a game may only pay off many steps later, which is what separates RL from simply learning from labels.

Start with linear regression

Linear regression (y = mx + c, just with more dimensions) is usually the first thing anyone implements, and it's still a strong first choice when the relationship between inputs and output is roughly linear. Try it before reaching for a neural network. If the linear model fits well, it is also far easier to explain to a stakeholder or auditor.

Here it is on scikit-learn's bundled diabetes dataset, so the snippet runs as-is with no download:

from sklearn.datasets import load_diabetes
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split

# 442 patients, 10 numeric features, target = disease progression after one year
X, y = load_diabetes(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

model = LinearRegression()
model.fit(X_train, y_train)
print(f"R^2 on test set: {model.score(X_test, y_test):.3f}")

Output:

R^2 on test set: 0.453

An R2 of 0.453 means the model explains about 45% of the variance in the held-out 20% of patients. That is modest, and it is exactly the kind of baseline number to beat before deciding a more complex model is worth its cost.

Decision trees split data on a sequence of yes/no questions and are worth using because a shallow tree's logic can be read line by line afterward. That stops being true once a tree grows hundreds of nodes deep, but it still beats most deep learning models, and it matters a lot when a decision needs to be explained to someone outside the modeling team. Neural networks (layers of connected nodes, loosely inspired by how neurons fire) are what you reach for once the pattern is too complex for the simpler methods to capture, but they need more data and you can no longer read off why they made a given prediction. Support vector machines find the boundary with the widest margin between classes and tend to hold up well on smaller, cleaner datasets where deep learning is overkill and the extra data a neural network would need simply isn't available.

MethodReach for it whenMain downside
Linear regressionRelationship looks roughly linear; explainability mattersUnderfits genuinely non-linear patterns
Decision treesA readable, auditable decision path is requiredProne to overfitting without pruning or ensembling
Neural networksPattern is too complex for simpler methods, enough data existsData-hungry, hard to interpret, easy to overfit on small datasets
Support vector machinesSmaller, cleaner dataset with a clear class boundaryKernel SVMs scale poorly to very large datasets

Biased data and personal data

A model trained on biased historical data reproduces that bias at scale, across far more decisions than any individual human could make. A lending model trained on past approval data can encode decades of discriminatory lending patterns without anyone writing a single explicitly discriminatory rule. The bias comes in through the historical outcomes the model is trained to reproduce. Amazon hit this with an experimental recruiting model: Reuters reported in 2018 that it had learned to downgrade resumes containing the word "women's", because it was trained on a decade of mostly male hiring, and the project was scrapped. This is the same failure mode covered in more depth in explainable AI for regulated industries, worth reading if this is a live concern for a specific deployment.

Data privacy is the other recurring problem: training data has to come from somewhere, and it is often personal information that people handed over for some other purpose. In March 2023 Italy's data protection authority temporarily blocked ChatGPT, citing among other things the lack of a legal basis for collecting personal data to train it. For your own projects, check which data you're allowed to train on, and report error rates per subgroup as well as overall, since a good average can hide a group the model fails.

Courses and practice

Andrew Ng's Coursera course is still a common starting point for the theory. The original 2011 course was replaced in 2022 by the Machine Learning Specialization, which teaches the same ideas in Python instead of Octave/MATLAB. Once the concepts click, Kaggle gives you real datasets and public notebooks to practice on. Stack Overflow is where most people end up anyway the first time scikit-learn throws an error they don't recognize.

Frequently asked questions

What's the main difference between supervised and unsupervised learning?
Supervised learning uses labeled examples to learn to predict a label for new data. Unsupervised learning finds structure in unlabeled data on its own, with no single correct answer to check the output against.

Which machine learning algorithm should I try first?
Linear regression, if the relationship between inputs and output looks roughly linear. It's the simplest to implement and the easiest to explain to someone outside the modeling team. Only move to a neural network once simpler methods can't capture the pattern and enough data exists to train one.

Is reinforcement learning the same as supervised learning?
No. Reinforcement learning has an agent take actions in an environment and learn from reward signals over time, rather than learning from a fixed labeled dataset. It usually needs far more samples and compute as a result, because the agent has to try actions to learn what works.

Where do the real-world risks in machine learning actually show up?
Mainly in biased training data, where a model reproduces historical bias at scale, and in data privacy, since training data is often personal information that people handed over for some other purpose.