Top 10 Machine Learning Algorithms: When to Use Each One (With Code)
Choosing the wrong algorithm wastes days of tuning. This guide cuts straight to: what problem each algorithm solves, a minimal working Python example, and when you should reach for something else instead. Every output shown below comes from running the code exactly as printed, with fixed random seeds.
All examples use scikit-learn unless noted. Install dependencies:
pip install scikit-learn xgboost
1. Linear Regression
Problem it solves: Predict a continuous value when the relationship between features and target is approximately linear.
from sklearn.linear_model import LinearRegression
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
import numpy as np
data = fetch_california_housing()
X, y = data.data, data.target
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = LinearRegression()
model.fit(X_train, y_train)
preds = model.predict(X_test)
print(f"RMSE: {np.sqrt(mean_squared_error(y_test, preds)):.3f}")
# Coefficient sign = direction of the effect. Features are on different
# scales, so compare magnitudes only after standardizing.
for name, coef in zip(data.feature_names, model.coef_):
print(f" {name}: {coef:+.4f}")
Output:
RMSE: 0.746
MedInc: +0.4487
HouseAge: +0.0097
AveRooms: -0.1233
AveBedrms: +0.7831
Population: -0.0000
AveOccup: -0.0035
Latitude: -0.4198
Longitude: -0.4337
AveBedrms has the largest raw coefficient only because its values span a narrow range (a one-unit change is a big change). Scale the features with StandardScaler before reading coefficient size as importance.
When NOT to use it: When features interact non-linearly, when you have many irrelevant features (use Ridge/Lasso instead), or when outliers dominate the loss. Non-normal residuals do not break the predictions; they matter for p-values and confidence intervals if you are doing statistical inference.
2. Logistic Regression
Problem it solves: Binary or multi-class classification with interpretable probability outputs. Your baseline for any classification problem.
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = LogisticRegression(max_iter=10000)
model.fit(X_train, y_train)
print(classification_report(y_test, model.predict(X_test)))
# Class probabilities (column 1 = probability of class 1)
probs = model.predict_proba(X_test)[:, 1]
Output:
precision recall f1-score support
0 0.97 0.91 0.94 43
1 0.95 0.99 0.97 71
accuracy 0.96 114
macro avg 0.96 0.95 0.95 114
weighted avg 0.96 0.96 0.96 114
When NOT to use it: When decision boundaries are highly non-linear. Try it first anyway: it's fast and gives you a baseline to beat.
3. Decision Tree
Problem it solves: Classification or regression with non-linear boundaries. Fully interpretable: you can print the exact rules it learned.
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
model = DecisionTreeClassifier(max_depth=3, random_state=42)
model.fit(X, y)
# Print the actual learned rules
print(export_text(model, feature_names=load_iris().feature_names))
Output:
|--- petal length (cm) <= 2.45
| |--- class: 0
|--- petal length (cm) > 2.45
| |--- petal width (cm) <= 1.75
| | |--- petal length (cm) <= 4.95
| | | |--- class: 1
| | |--- petal length (cm) > 4.95
| | | |--- class: 2
| |--- petal width (cm) > 1.75
| | |--- petal length (cm) <= 4.85
| | | |--- class: 2
| | |--- petal length (cm) > 4.85
| | | |--- class: 2
When NOT to use it: An unconstrained tree overfits noisy data easily, so limit max_depth or min_samples_leaf. Ensembles such as Random Forest or Gradient Boosting are usually more accurate; a single tree makes sense when a readable explanation matters more than the last few points of accuracy.
4. Random Forest
Problem it solves: Robust classification and regression by averaging many decorrelated trees. Needs no feature scaling, accepts NaN values directly in current scikit-learn (checked on 1.9.1), and gives feature importances out of the box. Categorical features still need encoding.
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
import numpy as np
X, y = load_breast_cancer(return_X_y=True)
model = RandomForestClassifier(n_estimators=200, max_features="sqrt", random_state=42, n_jobs=-1)
scores = cross_val_score(model, X, y, cv=5, scoring="roc_auc")
print(f"ROC-AUC: {scores.mean():.3f} ± {scores.std():.3f}")
# Feature importances
model.fit(X, y)
importances = sorted(zip(load_breast_cancer().feature_names, model.feature_importances_),
key=lambda x: -x[1])
for name, imp in importances[:5]:
print(f" {name}: {imp:.3f}")
Output:
ROC-AUC: 0.992 ± 0.006
worst perimeter: 0.143
worst area: 0.128
worst concave points: 0.119
mean concave points: 0.102
worst radius: 0.076
When NOT to use it: When you need a model you can explain to a non-technical stakeholder rule-by-rule. Large forests also take a lot of memory and are slower at prediction time; reduce n_estimators or max_depth if latency matters.
5. Support Vector Machine (SVM)
Problem it solves: High-accuracy classification, especially effective on high-dimensional data (text, images) and small-to-medium datasets where the margin between classes matters.
from sklearn.svm import SVC
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import accuracy_score
X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# SVM requires feature scaling
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
model = SVC(kernel="rbf", C=10, gamma="scale")
model.fit(X_train, y_train)
print(f"Accuracy: {accuracy_score(y_test, model.predict(X_test)):.3f}")
Output:
Accuracy: 0.981
When NOT to use it: Large datasets. Kernel SVC training time grows between quadratically and cubically with the number of samples, so it becomes impractical somewhere in the tens to hundreds of thousands of rows. For large datasets, use LinearSVC or SGDClassifier(loss="hinge") (both linear SVMs), or switch to gradient boosting.
6. K-Nearest Neighbors (KNN)
Problem it solves: Instance-based classification/regression with almost no training cost: fit just stores the data (or builds a search tree). Useful for recommendation-style problems where "similar inputs have similar outputs" is a safe assumption.
from sklearn.neighbors import KNeighborsClassifier
from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import accuracy_score
X, y = load_wine(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
scaler = StandardScaler() # KNN is distance-based, so scale features first
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
model = KNeighborsClassifier(n_neighbors=5)
model.fit(X_train, y_train)
print(f"Accuracy: {accuracy_score(y_test, model.predict(X_test)):.3f}")
Output:
Accuracy: 0.944
When NOT to use it: High-dimensional data (curse of dimensionality), large datasets (brute-force prediction compares every query with all n training points), or when you need probability calibration. It also has no built-in feature selection.
7. K-Means Clustering
Problem it solves: Unsupervised grouping of data into k clusters. Customer segmentation is the textbook use.
from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs
from sklearn.metrics import silhouette_score
import numpy as np
X, _ = make_blobs(n_samples=500, centers=4, random_state=42)
# Find optimal k using silhouette score
for k in range(2, 8):
km = KMeans(n_clusters=k, random_state=42, n_init=10)
labels = km.fit_predict(X)
score = silhouette_score(X, labels)
print(f"k={k} silhouette={score:.3f}")
# Fit with optimal k
best_model = KMeans(n_clusters=4, random_state=42, n_init=10)
X_labeled = best_model.fit_predict(X)
Output:
k=2 silhouette=0.596
k=3 silhouette=0.761
k=4 silhouette=0.791
k=5 silhouette=0.663
k=6 silhouette=0.561
k=7 silhouette=0.441
The silhouette score peaks at k=4, which matches the 4 centers the data was generated with.
When NOT to use it: Non-spherical clusters (try DBSCAN) or clusters with very different sizes and densities. K-Means also needs k up front; a silhouette loop like the one above or the elbow method helps choose it, but on real data the peak is rarely this clear.
8. Naive Bayes
Problem it solves: Fast probabilistic text classification. Despite the "naive" independence assumption, the fit call below took about 2 ms on 1,391 training documents in our run, and it is hard to beat as a first text baseline, especially with little labeled data.
from sklearn.naive_bayes import MultinomialNB
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report
# Example: 20 newsgroups text classification
from sklearn.datasets import fetch_20newsgroups
cats = ["sci.space", "rec.sport.hockey", "talk.politics.guns"]
data = fetch_20newsgroups(categories=cats, remove=("headers", "footers", "quotes"))
X_train, X_test, y_train, y_test = train_test_split(
data.data, data.target, test_size=0.2, random_state=42
)
vec = TfidfVectorizer(max_features=10000)
X_train_vec = vec.fit_transform(X_train)
X_test_vec = vec.transform(X_test)
model = MultinomialNB(alpha=0.1)
model.fit(X_train_vec, y_train)
print(classification_report(y_test, model.predict(X_test_vec), target_names=data.target_names))
Output:
precision recall f1-score support
rec.sport.hockey 0.94 0.93 0.94 125
sci.space 0.91 0.91 0.91 117
talk.politics.guns 0.92 0.94 0.93 106
accuracy 0.93 348
macro avg 0.92 0.93 0.93 348
weighted avg 0.93 0.93 0.93 348
Pass data.target_names, not cats, to the report. fetch_20newsgroups sorts the categories alphabetically, so label 0 is rec.sport.hockey. Passing cats runs without an error but prints the first two rows under the wrong names.
When NOT to use it: When features are strongly correlated (the independence assumption breaks down badly) or when you need well-calibrated probabilities beyond simple classification.
9. Gradient Boosting (XGBoost)
Problem it solves: Tabular data classification and regression. Gradient-boosted trees (XGBoost, LightGBM, CatBoost) appear in a large share of winning solutions for tabular Kaggle competitions. It builds trees sequentially, each correcting the errors of the previous one.
import xgboost as xgb
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
import numpy as np
X, y = fetch_california_housing(return_X_y=True)
# Hold out the test set FIRST and never let early stopping see it. A val
# set (carved out of the remaining training data) picks the stopping
# iteration, so the final test score is a genuine out-of-sample estimate.
X_train_full, X_test, y_train_full, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
X_train, X_val, y_train, y_val = train_test_split(X_train_full, y_train_full, test_size=0.2, random_state=42)
model = xgb.XGBRegressor(
n_estimators=2000, # upper bound; early stopping picks the real number
learning_rate=0.05,
max_depth=6,
subsample=0.8,
colsample_bytree=0.8,
early_stopping_rounds=20,
eval_metric="rmse",
random_state=42,
)
model.fit(X_train, y_train, eval_set=[(X_val, y_val)], verbose=False)
preds = model.predict(X_test)
print(f"RMSE: {np.sqrt(mean_squared_error(y_test, preds)):.3f}")
print(f"Best iteration: {model.best_iteration}")
Output:
RMSE: 0.444
Best iteration: 695
Compare this with 0.746 for linear regression on the same test split. Set n_estimators high enough that early stopping actually triggers: with n_estimators=500 this run hit the cap (best iteration 499) and scored 0.449.
When NOT to use it: Unstructured data (images, text, audio), where neural networks are the better fit. It also has more hyperparameters to tune than Random Forest; LightGBM and scikit-learn's HistGradientBoostingRegressor are worth benchmarking alongside it.
10. Neural Networks (MLP)
Problem it solves: Learning non-linear mappings from input to output. MLPClassifier is a plain fully connected network that trains on the CPU; the state of the art in images and text comes from CNNs and transformers built in PyTorch or TensorFlow, not from this class.
from sklearn.neural_network import MLPClassifier
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import classification_report
X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
model = MLPClassifier(
hidden_layer_sizes=(256, 128),
activation="relu",
max_iter=500,
early_stopping=True,
random_state=42,
)
model.fit(X_train, y_train)
print(classification_report(y_test, model.predict(X_test)))
Output:
precision recall f1-score support
0 1.00 1.00 1.00 33
1 1.00 0.96 0.98 28
2 0.97 1.00 0.99 33
3 1.00 0.94 0.97 34
4 1.00 1.00 1.00 46
5 0.92 0.94 0.93 47
6 0.97 0.97 0.97 35
7 0.97 0.97 0.97 34
8 0.97 0.97 0.97 30
9 0.93 0.95 0.94 40
accuracy 0.97 360
macro avg 0.97 0.97 0.97 360
weighted avg 0.97 0.97 0.97 360
When NOT to use it: Small tabular datasets, where simpler models usually match it with less tuning. On this 1,797-sample digits set the MLP reached 0.97 accuracy, while the SVM in section 5 reached 0.981 on the same split. For tabular data, try XGBoost first.
Quick Reference: Which Algorithm for Which Problem?
| Task | Start with | If that's not enough |
|---|---|---|
| Regression | Linear Regression | Random Forest, XGBoost |
| Binary classification | Logistic Regression | XGBoost, SVM |
| Multi-class classification | Logistic Regression | Random Forest, Neural Net |
| Text classification | Naive Bayes + TF-IDF | Fine-tuned BERT |
| Clustering | K-Means | DBSCAN, Gaussian Mixture |
| Image / audio | CNN (PyTorch/TF) | Fine-tune pretrained model |