Skip to content

Top 10 Machine Learning Algorithms: When to Use Each One (With Code)

Tested with: Python 3.12.3, scikit-learn 1.9.1, XGBoost 3.4.1, NumPy 2.5.3, SciPy 1.18.1 (CPU, Ubuntu 24.04). Last run 2026-09-27.

Choosing the wrong algorithm wastes days of tuning. This guide cuts straight to: what problem each algorithm solves, a minimal working Python example, and when you should reach for something else instead. Every output shown below comes from running the code exactly as printed, with fixed random seeds.

All examples use scikit-learn unless noted. Install dependencies:

pip install scikit-learn xgboost

1. Linear Regression

Problem it solves: Predict a continuous value when the relationship between features and target is approximately linear.

from sklearn.linear_model import LinearRegression
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
import numpy as np

data = fetch_california_housing()
X, y = data.data, data.target
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

model = LinearRegression()
model.fit(X_train, y_train)
preds = model.predict(X_test)
print(f"RMSE: {np.sqrt(mean_squared_error(y_test, preds)):.3f}")
# Coefficient sign = direction of the effect. Features are on different
# scales, so compare magnitudes only after standardizing.
for name, coef in zip(data.feature_names, model.coef_):
    print(f"  {name}: {coef:+.4f}")

Output:

RMSE: 0.746
  MedInc: +0.4487
  HouseAge: +0.0097
  AveRooms: -0.1233
  AveBedrms: +0.7831
  Population: -0.0000
  AveOccup: -0.0035
  Latitude: -0.4198
  Longitude: -0.4337

AveBedrms has the largest raw coefficient only because its values span a narrow range (a one-unit change is a big change). Scale the features with StandardScaler before reading coefficient size as importance.

When NOT to use it: When features interact non-linearly, when you have many irrelevant features (use Ridge/Lasso instead), or when outliers dominate the loss. Non-normal residuals do not break the predictions; they matter for p-values and confidence intervals if you are doing statistical inference.


2. Logistic Regression

Problem it solves: Binary or multi-class classification with interpretable probability outputs. Your baseline for any classification problem.

from sklearn.linear_model import LogisticRegression
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

model = LogisticRegression(max_iter=10000)
model.fit(X_train, y_train)
print(classification_report(y_test, model.predict(X_test)))
# Class probabilities (column 1 = probability of class 1)
probs = model.predict_proba(X_test)[:, 1]

Output:

              precision    recall  f1-score   support

           0       0.97      0.91      0.94        43
           1       0.95      0.99      0.97        71

    accuracy                           0.96       114
   macro avg       0.96      0.95      0.95       114
weighted avg       0.96      0.96      0.96       114

When NOT to use it: When decision boundaries are highly non-linear. Try it first anyway: it's fast and gives you a baseline to beat.


3. Decision Tree

Problem it solves: Classification or regression with non-linear boundaries. Fully interpretable: you can print the exact rules it learned.

from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)
model = DecisionTreeClassifier(max_depth=3, random_state=42)
model.fit(X, y)

# Print the actual learned rules
print(export_text(model, feature_names=load_iris().feature_names))

Output:

|--- petal length (cm) <= 2.45
|   |--- class: 0
|--- petal length (cm) >  2.45
|   |--- petal width (cm) <= 1.75
|   |   |--- petal length (cm) <= 4.95
|   |   |   |--- class: 1
|   |   |--- petal length (cm) >  4.95
|   |   |   |--- class: 2
|   |--- petal width (cm) >  1.75
|   |   |--- petal length (cm) <= 4.85
|   |   |   |--- class: 2
|   |   |--- petal length (cm) >  4.85
|   |   |   |--- class: 2

When NOT to use it: An unconstrained tree overfits noisy data easily, so limit max_depth or min_samples_leaf. Ensembles such as Random Forest or Gradient Boosting are usually more accurate; a single tree makes sense when a readable explanation matters more than the last few points of accuracy.


4. Random Forest

Problem it solves: Robust classification and regression by averaging many decorrelated trees. Needs no feature scaling, accepts NaN values directly in current scikit-learn (checked on 1.9.1), and gives feature importances out of the box. Categorical features still need encoding.

from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = load_breast_cancer(return_X_y=True)
model = RandomForestClassifier(n_estimators=200, max_features="sqrt", random_state=42, n_jobs=-1)
scores = cross_val_score(model, X, y, cv=5, scoring="roc_auc")
print(f"ROC-AUC: {scores.mean():.3f} ± {scores.std():.3f}")

# Feature importances
model.fit(X, y)
importances = sorted(zip(load_breast_cancer().feature_names, model.feature_importances_),
                     key=lambda x: -x[1])
for name, imp in importances[:5]:
    print(f"  {name}: {imp:.3f}")

Output:

ROC-AUC: 0.992 ± 0.006
  worst perimeter: 0.143
  worst area: 0.128
  worst concave points: 0.119
  mean concave points: 0.102
  worst radius: 0.076

When NOT to use it: When you need a model you can explain to a non-technical stakeholder rule-by-rule. Large forests also take a lot of memory and are slower at prediction time; reduce n_estimators or max_depth if latency matters.


5. Support Vector Machine (SVM)

Problem it solves: High-accuracy classification, especially effective on high-dimensional data (text, images) and small-to-medium datasets where the margin between classes matters.

from sklearn.svm import SVC
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import accuracy_score

X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# SVM requires feature scaling
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)

model = SVC(kernel="rbf", C=10, gamma="scale")
model.fit(X_train, y_train)
print(f"Accuracy: {accuracy_score(y_test, model.predict(X_test)):.3f}")

Output:

Accuracy: 0.981

When NOT to use it: Large datasets. Kernel SVC training time grows between quadratically and cubically with the number of samples, so it becomes impractical somewhere in the tens to hundreds of thousands of rows. For large datasets, use LinearSVC or SGDClassifier(loss="hinge") (both linear SVMs), or switch to gradient boosting.


6. K-Nearest Neighbors (KNN)

Problem it solves: Instance-based classification/regression with almost no training cost: fit just stores the data (or builds a search tree). Useful for recommendation-style problems where "similar inputs have similar outputs" is a safe assumption.

from sklearn.neighbors import KNeighborsClassifier
from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import accuracy_score

X, y = load_wine(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

scaler = StandardScaler()  # KNN is distance-based, so scale features first
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)

model = KNeighborsClassifier(n_neighbors=5)
model.fit(X_train, y_train)
print(f"Accuracy: {accuracy_score(y_test, model.predict(X_test)):.3f}")

Output:

Accuracy: 0.944

When NOT to use it: High-dimensional data (curse of dimensionality), large datasets (brute-force prediction compares every query with all n training points), or when you need probability calibration. It also has no built-in feature selection.


7. K-Means Clustering

Problem it solves: Unsupervised grouping of data into k clusters. Customer segmentation is the textbook use.

from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs
from sklearn.metrics import silhouette_score
import numpy as np

X, _ = make_blobs(n_samples=500, centers=4, random_state=42)

# Find optimal k using silhouette score
for k in range(2, 8):
    km = KMeans(n_clusters=k, random_state=42, n_init=10)
    labels = km.fit_predict(X)
    score = silhouette_score(X, labels)
    print(f"k={k}  silhouette={score:.3f}")

# Fit with optimal k
best_model = KMeans(n_clusters=4, random_state=42, n_init=10)
X_labeled = best_model.fit_predict(X)

Output:

k=2  silhouette=0.596
k=3  silhouette=0.761
k=4  silhouette=0.791
k=5  silhouette=0.663
k=6  silhouette=0.561
k=7  silhouette=0.441

The silhouette score peaks at k=4, which matches the 4 centers the data was generated with.

When NOT to use it: Non-spherical clusters (try DBSCAN) or clusters with very different sizes and densities. K-Means also needs k up front; a silhouette loop like the one above or the elbow method helps choose it, but on real data the peak is rarely this clear.


8. Naive Bayes

Problem it solves: Fast probabilistic text classification. Despite the "naive" independence assumption, the fit call below took about 2 ms on 1,391 training documents in our run, and it is hard to beat as a first text baseline, especially with little labeled data.

from sklearn.naive_bayes import MultinomialNB
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report

# Example: 20 newsgroups text classification
from sklearn.datasets import fetch_20newsgroups
cats = ["sci.space", "rec.sport.hockey", "talk.politics.guns"]
data = fetch_20newsgroups(categories=cats, remove=("headers", "footers", "quotes"))

X_train, X_test, y_train, y_test = train_test_split(
    data.data, data.target, test_size=0.2, random_state=42
)
vec = TfidfVectorizer(max_features=10000)
X_train_vec = vec.fit_transform(X_train)
X_test_vec = vec.transform(X_test)

model = MultinomialNB(alpha=0.1)
model.fit(X_train_vec, y_train)
print(classification_report(y_test, model.predict(X_test_vec), target_names=data.target_names))

Output:

                    precision    recall  f1-score   support

  rec.sport.hockey       0.94      0.93      0.94       125
         sci.space       0.91      0.91      0.91       117
talk.politics.guns       0.92      0.94      0.93       106

          accuracy                           0.93       348
         macro avg       0.92      0.93      0.93       348
      weighted avg       0.93      0.93      0.93       348

Pass data.target_names, not cats, to the report. fetch_20newsgroups sorts the categories alphabetically, so label 0 is rec.sport.hockey. Passing cats runs without an error but prints the first two rows under the wrong names.

When NOT to use it: When features are strongly correlated (the independence assumption breaks down badly) or when you need well-calibrated probabilities beyond simple classification.


9. Gradient Boosting (XGBoost)

Problem it solves: Tabular data classification and regression. Gradient-boosted trees (XGBoost, LightGBM, CatBoost) appear in a large share of winning solutions for tabular Kaggle competitions. It builds trees sequentially, each correcting the errors of the previous one.

import xgboost as xgb
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error
import numpy as np

X, y = fetch_california_housing(return_X_y=True)
# Hold out the test set FIRST and never let early stopping see it. A val
# set (carved out of the remaining training data) picks the stopping
# iteration, so the final test score is a genuine out-of-sample estimate.
X_train_full, X_test, y_train_full, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
X_train, X_val, y_train, y_val = train_test_split(X_train_full, y_train_full, test_size=0.2, random_state=42)

model = xgb.XGBRegressor(
    n_estimators=2000,  # upper bound; early stopping picks the real number
    learning_rate=0.05,
    max_depth=6,
    subsample=0.8,
    colsample_bytree=0.8,
    early_stopping_rounds=20,
    eval_metric="rmse",
    random_state=42,
)
model.fit(X_train, y_train, eval_set=[(X_val, y_val)], verbose=False)
preds = model.predict(X_test)
print(f"RMSE: {np.sqrt(mean_squared_error(y_test, preds)):.3f}")
print(f"Best iteration: {model.best_iteration}")

Output:

RMSE: 0.444
Best iteration: 695

Compare this with 0.746 for linear regression on the same test split. Set n_estimators high enough that early stopping actually triggers: with n_estimators=500 this run hit the cap (best iteration 499) and scored 0.449.

When NOT to use it: Unstructured data (images, text, audio), where neural networks are the better fit. It also has more hyperparameters to tune than Random Forest; LightGBM and scikit-learn's HistGradientBoostingRegressor are worth benchmarking alongside it.


10. Neural Networks (MLP)

Problem it solves: Learning non-linear mappings from input to output. MLPClassifier is a plain fully connected network that trains on the CPU; the state of the art in images and text comes from CNNs and transformers built in PyTorch or TensorFlow, not from this class.

from sklearn.neural_network import MLPClassifier
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import classification_report

X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)

model = MLPClassifier(
    hidden_layer_sizes=(256, 128),
    activation="relu",
    max_iter=500,
    early_stopping=True,
    random_state=42,
)
model.fit(X_train, y_train)
print(classification_report(y_test, model.predict(X_test)))

Output:

              precision    recall  f1-score   support

           0       1.00      1.00      1.00        33
           1       1.00      0.96      0.98        28
           2       0.97      1.00      0.99        33
           3       1.00      0.94      0.97        34
           4       1.00      1.00      1.00        46
           5       0.92      0.94      0.93        47
           6       0.97      0.97      0.97        35
           7       0.97      0.97      0.97        34
           8       0.97      0.97      0.97        30
           9       0.93      0.95      0.94        40

    accuracy                           0.97       360
   macro avg       0.97      0.97      0.97       360
weighted avg       0.97      0.97      0.97       360

When NOT to use it: Small tabular datasets, where simpler models usually match it with less tuning. On this 1,797-sample digits set the MLP reached 0.97 accuracy, while the SVM in section 5 reached 0.981 on the same split. For tabular data, try XGBoost first.


Quick Reference: Which Algorithm for Which Problem?

TaskStart withIf that's not enough
RegressionLinear RegressionRandom Forest, XGBoost
Binary classificationLogistic RegressionXGBoost, SVM
Multi-class classificationLogistic RegressionRandom Forest, Neural Net
Text classificationNaive Bayes + TF-IDFFine-tuned BERT
ClusteringK-MeansDBSCAN, Gaussian Mixture
Image / audioCNN (PyTorch/TF)Fine-tune pretrained model