Skip to content

Fix sklearn NotFittedError: "This Estimator Is Not Fitted Yet"

Tested with: scikit-learn 1.9.1, numpy 2.5.3, scipy 1.18.1, joblib 1.6.0, pandas 3.0.6, Python 3.12.3, Linux (Intel i5-7500). The same script also ran on scikit-learn 1.3.2 with numpy 1.26.4; the two differences are noted below. Last run 2026-09-27.

TL;DR: the object you called predict(), transform() or score() on has never been fitted. Usually you fitted a different object than the one you are using.
  1. Plain case: call fit(X_train, y_train) on that exact object first.
  2. After cross_val_score, clone() or GridSearchCV: these fit copies and leave your original untouched. Fit the original yourself, or use grid.best_estimator_ / grid.predict().
  3. Pipelines: use pipe.predict() or pipe.named_steps['name'], not a separate transformer you built next to it.
  4. Loaded from disk: the file was written before fit(). Call check_is_fitted(model) right before joblib.dump().
  5. Your own estimator: store learned values in attributes ending in _ (for example self.majority_), or define __sklearn_is_fitted__.

sklearn.exceptions.NotFittedError is raised when a method that needs learned parameters (predict(), transform(), score(), predict_proba()) runs on an estimator object that has never been fitted. The cases below differ only in how you ended up holding an unfitted object.

Reproduced: Each Case and the Fix That Worked

Every row was run on the versions in the header, on the iris dataset. The error text is copied from the output. In 1.9.1 every message has the same shape: This <ClassName> instance is not fitted yet. Call 'fit' with appropriate arguments before using this estimator.

What was runExact error / outputFix that worked
RandomForestClassifier().predict(X), no fitNotFittedError: This RandomForestClassifier instance is not fitted yet. ...clf.fit(X_train, y_train) first (test accuracy 1.0)
StandardScaler().transform(X_train), PolynomialFeatures().transform(...)Same message with StandardScaler / PolynomialFeaturesfit_transform on train, transform on test
pipe.fit(...), then transform() on a separate StandardScaler()Same message with StandardScalerpipe.named_steps['scaler'] or pipe.predict()
pipe.predict(X) before pipe.fit1.9.1: This Pipeline instance is not fitted yet.
1.3.2: This StandardScaler instance is not fitted yet. (first step)
pipe.fit(X_train, y_train)
clone(fitted_model).predict(X)Same message with LogisticRegressioncopy.deepcopy(fitted_model) keeps the fit; clone is meant to drop it
cross_val_score(model, X, y, cv=5), then model.predict(X_test)Scores [0.967 1. 0.933 0.967 1.], then the same error with LogisticRegressionmodel.fit(X, y) after CV, or cross_validate(..., return_estimator=True)
GridSearchCV(base, ...).fit(...), then base.predict(X_test)Same message with SVCgrid.best_estimator_ or grid.predict()
GridSearchCV(..., refit=False), then grid.predict / grid.best_estimator_Not NotFittedError. 1.9.1: AttributeError: This 'GridSearchCV' has no attribute 'predict' and 'GridSearchCV' object has no attribute 'best_estimator_'Keep refit=True, or fit a new model with grid.best_params_
pickle / joblib dump of an unfitted model, load, predictSame message with LogisticRegression / GradientBoostingClassifierFit before dump; the fitted round trip predicted [1, 0, 2]
check_is_fitted(rf, attributes=['n_estimators_', 'estimators_']) on a fitted forestNotFittedError (there is no n_estimators_, and all listed names must exist)attributes=['estimators_'], or no attributes at all
check_is_fitted(RandomForestClassifier) (the class)TypeError: <class 'sklearn.ensemble._forest.RandomForestClassifier'> is a class, not an instance.Pass an instance
Custom estimator that stores self.majority (no trailing _) and calls check_is_fitted(self)This NoUnderscore instance is not fitted yet. even right after fit()Rename to self.majority_, or add __sklearn_is_fitted__

What the Error Means

Every scikit-learn estimator follows the same contract: call fit(X, y) first, then call predict(), transform(), score(), etc. Internally, fit() stores learned attributes (weights, feature names, vocabulary, scaling parameters) on the estimator instance. Methods like predict() look for those attributes. If they are absent, scikit-learn raises NotFittedError.

Here is the full traceback for predict() on an unfitted random forest:

Traceback (most recent call last):
  File "repro.py", line 30, in <module>
    RandomForestClassifier(n_estimators=100, random_state=42).predict(X)
  File ".../site-packages/sklearn/ensemble/_forest.py", line 903, in predict
    proba = self.predict_proba(X)
            ^^^^^^^^^^^^^^^^^^^^^
  File ".../site-packages/sklearn/ensemble/_forest.py", line 943, in predict_proba
    check_is_fitted(self)
  File ".../site-packages/sklearn/utils/validation.py", line 1721, in check_is_fitted
    raise NotFittedError(msg % {"name": type(estimator).__name__})
sklearn.exceptions.NotFittedError: This RandomForestClassifier instance is not fitted yet. Call 'fit' with appropriate arguments before using this estimator.

That is the real traceback from scikit-learn 1.9.1 (site-packages path shortened). Line numbers move between releases; in 1.3.2 the same frames were at lines 823, 863 and 1461.

The key phrase is "Call 'fit' with appropriate arguments before using this estimator." The class name in the message tells you which object is unfitted; search your code for where that object was created and check whether fit() was ever called on that same variable.

Note: The error lives at sklearn.exceptions.NotFittedError. It is also importable from sklearn.utils.validation (checked: it is the same class), so either import works in except NotFittedError. It subclasses both ValueError and AttributeError, so a broad except ValueError will also swallow it.

Cause 1: predict() Called Before fit()

You created the estimator and used it without training it first.

# WRONG: predict() called on a fresh, unfitted instance
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)

clf = RandomForestClassifier(n_estimators=100, random_state=42)
# Missing: clf.fit(X_train, y_train)
predictions = clf.predict(X)  # <-- NotFittedError here
# CORRECT: call fit() before predict()
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

clf = RandomForestClassifier(n_estimators=100, random_state=42)
clf.fit(X_train, y_train)          # learn from training data first
predictions = clf.predict(X_test)

This pattern also applies to transformers. Calling scaler.transform(X) without a preceding scaler.fit(X) or scaler.fit_transform(X) produces the same error:

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
# WRONG: scaler.transform(X_train)
# CORRECT:
X_train_scaled = scaler.fit_transform(X_train)  # fits AND transforms in one call
X_test_scaled  = scaler.transform(X_test)        # uses parameters learned on train set
Common mistake: calling fit_transform() on the test set. This re-fits the scaler to test data, leaking information and invalidating your evaluation. Always fit on train, transform on test.

Cause 2: Fitting the Pipeline, Then Calling a Step Directly

When you use sklearn.pipeline.Pipeline, fitting the pipeline calls fit() on each step in sequence. The state is stored on the step objects inside the pipeline. A common mistake is to create a separate standalone estimator, fit the pipeline, and then call predict() or transform() on the standalone object, which was never fitted.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC

# Create a standalone scaler AND a pipeline that contains its OWN scaler
standalone_scaler = StandardScaler()

pipeline = Pipeline([
    ('scaler', StandardScaler()),   # this is a different object from standalone_scaler
    ('clf', SVC()),
])

pipeline.fit(X_train, y_train)

# WRONG: standalone_scaler was never fitted, NotFittedError
X_scaled = standalone_scaler.transform(X_test)
# CORRECT option A: access the fitted step from inside the pipeline
fitted_scaler = pipeline.named_steps['scaler']
X_scaled = fitted_scaler.transform(X_test)

# CORRECT option B: just use the pipeline for everything
predictions = pipeline.predict(X_test)   # pipeline handles transform internally

Calling pipe.predict() before pipe.fit() also raises, but the class name in the message depends on the version: 1.9.1 says This Pipeline instance is not fitted yet, while 1.3.2 named the first step (This StandardScaler instance ...). If the message names a transformer you never touched directly, check whether it is sitting inside an unfitted pipeline.

Cause 3: Pickle or joblib File Saved Before fit()

Here the error shows up in the script that loads the model, not in the one that saved it. Dumping an unfitted LogisticRegression with pickle (or an unfitted GradientBoostingClassifier with joblib) and loading it back gives the usual This LogisticRegression instance is not fitted yet. on the first predict().

import pickle
from sklearn.linear_model import LogisticRegression

# WRONG: model pickled before training
clf = LogisticRegression()
with open('model.pkl', 'wb') as f:
    pickle.dump(clf, f)   # unfitted

# ... later, in inference.py ...
with open('model.pkl', 'rb') as f:
    loaded_clf = pickle.load(f)

loaded_clf.predict(X_new)  # <-- NotFittedError: model was never fitted
import pickle
from sklearn.linear_model import LogisticRegression

# CORRECT: fit first, then serialize
clf = LogisticRegression()
clf.fit(X_train, y_train)   # <-- must happen before pickle

with open('model.pkl', 'wb') as f:
    pickle.dump(clf, f)

# Now the loaded model is safe to use
with open('model.pkl', 'rb') as f:
    loaded_clf = pickle.load(f)

predictions = loaded_clf.predict(X_new)

The same with joblib (the fitted round trip predicted [1, 0, 2] on the first three test rows):

import joblib
from sklearn.ensemble import GradientBoostingClassifier

clf = GradientBoostingClassifier()
clf.fit(X_train, y_train)

joblib.dump(clf, 'model.joblib')

# In inference
clf = joblib.load('model.joblib')
predictions = clf.predict(X_test)
Diagnostic tip: If you receive NotFittedError from a loaded model, check your training script's save logic. Call check_is_fitted(clf) immediately before the dump, so a broken training run fails there instead of in the service that loads the file.

Cause 4: Manual Step Order: transform() Before fit_transform()

Same error as the scaler case in Cause 1, but in a chain of transformers: after a refactor, one step gets transform() where it needed fit_transform(). The message names the step (This PolynomialFeatures instance is not fitted yet.).

from sklearn.preprocessing import StandardScaler, PolynomialFeatures
from sklearn.linear_model import Ridge

scaler = StandardScaler()
poly   = PolynomialFeatures(degree=2)
ridge  = Ridge(alpha=1.0)

# WRONG order: transform before fit
X_poly = poly.transform(X_train)           # NotFittedError: poly never fitted
X_scaled = scaler.fit_transform(X_poly)
ridge.fit(X_scaled, y_train)
# CORRECT order
X_poly   = poly.fit_transform(X_train)    # fit AND transform train set
X_scaled = scaler.fit_transform(X_poly)   # fit AND transform train set
ridge.fit(X_scaled, y_train)

# For test set: transform only, do NOT refit
X_test_poly   = poly.transform(X_test)
X_test_scaled = scaler.transform(X_test_poly)
predictions   = ridge.predict(X_test_scaled)

A Pipeline removes the ordering problem, since fit() runs fit_transform on each step in order:

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, PolynomialFeatures
from sklearn.linear_model import Ridge

pipe = Pipeline([
    ('poly',   PolynomialFeatures(degree=2)),
    ('scaler', StandardScaler()),
    ('model',  Ridge(alpha=1.0)),
])

pipe.fit(X_train, y_train)        # fits every step in order, automatically
predictions = pipe.predict(X_test) # transforms through every step, then predicts

Cause 5: clone() and cross_val_score Fit Copies, Not Your Model

sklearn.base.clone() copies an estimator's parameters and deliberately drops everything it learned. cross_val_score, cross_validate and the search classes all clone the estimator for each fold, so the object you passed in is still unfitted when they return.

from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score

model = LogisticRegression(max_iter=1000)
scores = cross_val_score(model, X, y, cv=5)
print(scores.round(3), scores.mean().round(3))
# [0.967 1.    0.933 0.967 1.   ] 0.973

model.predict(X_test)
# NotFittedError: This LogisticRegression instance is not fitted yet. ...
# FIX A: cross-validation is for estimating the score; fit the final model yourself
model.fit(X, y)
model.predict(X_test)

# FIX B: keep the per-fold models
from sklearn.model_selection import cross_validate
res = cross_validate(LogisticRegression(max_iter=1000), X, y, cv=5, return_estimator=True)
res["estimator"][0].predict(X_test[:3])   # array([1, 0, 2])

The same applies to copying a fitted model: clone(fitted).predict(X) raises NotFittedError, while copy.deepcopy(fitted) keeps the learned state and predicts normally.

Cause 6: GridSearchCV: the Estimator You Passed Stays Unfitted

GridSearchCV clones the estimator for every candidate and, with the default refit=True, fits one more clone on the full training set and stores it as best_estimator_. The object you passed in is kept as grid.estimator and is never fitted.

from sklearn.model_selection import GridSearchCV
from sklearn.svm import SVC

base = SVC()
grid = GridSearchCV(base, {"C": [0.1, 1, 10]}, cv=5).fit(X_train, y_train)

base.predict(X_test)
# NotFittedError: This SVC instance is not fitted yet. ...

grid.best_params_                        # {'C': 1}
grid.best_estimator_.score(X_test, y_test)  # 1.0
grid.predict(X_test)                     # same as best_estimator_.predict

With refit=False there is no refitted model, and the error is not NotFittedError. In 1.9.1, grid.predict raises AttributeError: This 'GridSearchCV' has no attribute 'predict' and grid.best_estimator_ raises AttributeError: 'GridSearchCV' object has no attribute 'best_estimator_'. (1.3.2 gave a longer message pointing at refit=False.) Use SVC(**grid.best_params_).fit(X_train, y_train) in that case. Calling predict on a GridSearchCV that was never fitted gives the normal This GridSearchCV instance is not fitted yet.

Cause 7: Custom Estimators Without Trailing-Underscore Attributes

With no attributes argument, check_is_fitted() looks for instance attributes whose names end in _ (and do not start with __). If your own estimator stores what it learned under a plain name, check_is_fitted(self) fails even after a successful fit():

import numpy as np
from sklearn.base import BaseEstimator, ClassifierMixin
from sklearn.utils.validation import check_is_fitted

class NoUnderscore(ClassifierMixin, BaseEstimator):
    def fit(self, X, y):
        self.majority = np.bincount(y).argmax()   # no trailing underscore
        return self
    def predict(self, X):
        check_is_fitted(self)
        return np.full(len(X), self.majority)

NoUnderscore().fit(X_train, y_train).predict(X_test)
# NotFittedError: This NoUnderscore instance is not fitted yet. ...

Two fixes, both verified:

# FIX A: follow the convention, learned state ends in "_"
def fit(self, X, y):
    self.majority_ = np.bincount(y).argmax()
    return self

# FIX B: tell scikit-learn explicitly
def __sklearn_is_fitted__(self):
    return getattr(self, "_fitted", False)   # set self._fitted = True in fit()

The opposite mistake is quieter: an estimator that sets self.model_ = None in __init__ passes check_is_fitted before it is ever fitted, so the guard never fires. Only create _-suffixed attributes inside fit().

Using check_is_fitted() Proactively

scikit-learn exposes the same internal check it uses in predict() and transform() as a public utility: sklearn.utils.validation.check_is_fitted(). Call it yourself where a model enters your code, for example right after joblib.load(), so the failure points at the file and not at the first request.

from sklearn.utils.validation import check_is_fitted
from sklearn.exceptions import NotFittedError
from sklearn.ensemble import RandomForestClassifier

clf = RandomForestClassifier()

# Guard before predict: fail with your own message
try:
    check_is_fitted(clf)
except NotFittedError:
    raise RuntimeError(
        "Model must be trained before inference. "
        "Call clf.fit(X_train, y_train) first."
    )

check_is_fitted() checks for the presence of any attributes ending in _ (the scikit-learn convention for fitted parameters, for example coef_, n_features_in_, classes_). Estimators that learn nothing, such as FunctionTransformer(), pass the check without being fitted. You can also check for specific attributes; every name you list must exist unless you pass all_or_any=any:

from sklearn.utils.validation import check_is_fitted

# Check for a specific fitted attribute (RandomForest sets estimators_;
# there is no n_estimators_, and listing it raises even on a fitted forest)
check_is_fitted(clf, attributes=['estimators_'])

# Custom message; %(name)s is replaced by the class name
check_is_fitted(clf, 'estimators_', msg="%(name)s: train it first")
# unfitted -> NotFittedError: RandomForestClassifier: train it first

# A boolean helper
def is_fitted(estimator):
    """Return True if the estimator has been fitted."""
    try:
        check_is_fitted(estimator)
        return True
    except NotFittedError:
        return False

clf = RandomForestClassifier()
print(is_fitted(clf))          # False

clf.fit(X_train, y_train)
print(is_fitted(clf))          # True

Build a safe predict wrapper

If a service loads a model file at startup, a guard turns the scikit-learn error into one that says what to check. On an unfitted forest it raised RuntimeError: RandomForestClassifier is not fitted. Ensure the model training pipeline ran successfully and the model artifact was saved after fit().

from sklearn.utils.validation import check_is_fitted
from sklearn.exceptions import NotFittedError

def safe_predict(estimator, X):
    """Predict with a clear error if the model is not ready."""
    try:
        check_is_fitted(estimator)
    except NotFittedError as exc:
        raise RuntimeError(
            f"{type(estimator).__name__} is not fitted. "
            "Ensure the model training pipeline ran successfully "
            "and the model artifact was saved after fit()."
        ) from exc
    return estimator.predict(X)

# Usage
predictions = safe_predict(clf, X_test)

Quick-reference: which attributes signal a fitted estimator?

Different estimators set different trailing-underscore attributes after fitting. Each line below returned True after fit() in 1.9.1:

# Classifiers
hasattr(clf, 'classes_')          # LogisticRegression, RandomForest, SVC, etc.
hasattr(clf, 'estimators_')       # RandomForest, GradientBoosting

# Regressors
hasattr(reg, 'coef_')             # LinearRegression, Ridge, Lasso
hasattr(reg, 'feature_importances_')  # tree-based regressors (False before fit)

# Transformers
hasattr(scaler, 'mean_')          # StandardScaler
hasattr(scaler, 'scale_')         # StandardScaler, RobustScaler
hasattr(enc, 'categories_')       # OneHotEncoder
hasattr(imputer, 'statistics_')   # SimpleImputer
hasattr(pca, 'components_')       # PCA