Skip to content

Fix sklearn ValueError: Input X Contains NaN

Tested with: scikit-learn 1.9.1, pandas 3.0.6, numpy 2.5.3, scipy 1.18.1, Python 3.12.3, Linux (Intel i5-7500). Tree checks also run on scikit-learn 1.3.2 and 1.5.2. Last run 2026-09-27.

TL;DR: scikit-learn found a missing value in your features (or in y) and the estimator can't handle it. Run df.isna().sum() to see which columns, then pick one:
  1. Impute inside a Pipeline: make_pipeline(SimpleImputer(strategy='median'), StandardScaler(), LogisticRegression()). Split first, then fit, so test rows never feed the imputer.
  2. Use an estimator that takes NaN: in 1.9.1, HistGradientBoosting*, DecisionTree*, RandomForest*, ExtraTree(s)* all fit on NaN without complaint. On 1.3.2, of DecisionTreeClassifier, RandomForestClassifier and ExtraTreesClassifier, only the single tree did.
  3. The "infinity" variant: df.replace([np.inf, -np.inf], np.nan), then impute. SimpleImputer itself rejects inf.
  4. NaN in y: drop those rows. No estimator accepts it and imputing a target is rarely what you want.

You call model.fit(X, y) and get this (exact text from scikit-learn 1.9.1, the second line is one long line):

ValueError: Input X contains NaN.
LogisticRegression does not accept missing values encoded as NaN natively. For supervised learning, you might want to consider sklearn.ensemble.HistGradientBoostingClassifier and Regressor which accept missing values encoded as NaNs natively. Alternatively, it is possible to preprocess the data, for instance by using an imputer transformer in a pipeline or drop samples with missing values. See https://scikit-learn.org/stable/modules/impute.html You can find a list of all estimators that handle NaN values at the following page: https://scikit-learn.org/stable/modules/impute.html#estimators-that-handle-nan-values

The estimator name changes (LinearRegression, SVC, KNeighborsClassifier, GradientBoostingClassifier all produced the same text with their own name), and so does the call: predict() on a row with NaN raises the identical message even if training data was clean. Dropping every row with a NaN is usually the worst fix. Below: how to find the NaNs, which imputer to use, and how to wire it so test data doesn't leak into training.

Reproduced: Each Error and the Fix That Worked

Every row was run on the versions in the header, error text copied from the output.

What was passedExact error / resultFix that worked
np.nan in X, LogisticRegression().fitValueError: Input X contains NaN. (plus the long hint above)SimpleImputer in a Pipeline
Nullable Int64 column with pd.NASame Input X contains NaN. messageSame
np.inf or -np.inf in X (e.g. a ratio divided by 0)ValueError: Input X contains infinity or a value too large for dtype('float64').replace([np.inf, -np.inf], np.nan), then impute
Finite 1e39 in X, RandomForestClassifierValueError: Input X contains infinity or a value too large for dtype('float32').Trees cast to float32 (max about 3.4e38): clip or rescale the column
np.inf into SimpleImputerValueError: Input X contains infinity or a value too large for dtype('float64').Replace inf before the imputer
NaN in y (LogisticRegression, LinearRegression, RandomForest, HistGradientBoosting)ValueError: Input y contains NaN.Drop rows where y is NaN
NaN in X, HistGradientBoostingClassifier, RandomForest*, DecisionTree*, ExtraTree(s)ClassifierFit and predict succeeded (1.9.1)None needed
StandardScaler().fit_transform on NaNNo error: NaN passed straight through to the outputPut the imputer before the scaler
from sklearn.impute import IterativeImputerImportError: IterativeImputer is experimental and the API might change without any deprecation cycle. To use it, you need to explicitly import enable_iterative_imputer:from sklearn.experimental import enable_iterative_imputer first
Old tutorial code: df.fillna(method='ffill') on pandas 3.0.6TypeError: NDFrame.fillna() got an unexpected keyword argument 'method'df.ffill()

Note that a very large but finite value like 1e300 did not raise with LogisticRegression: the "too large" part only bites when the estimator converts to a smaller dtype, as tree models do with float32.


Step 1: Find Which Columns Have NaN

The error tells you NaN exists but not where. Locate it first.

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'age': [25, np.nan, 35, 40, np.nan],
    'income': [50000, 60000, np.nan, 80000, 70000],
    'category': ['A', 'B', 'A', np.nan, 'B'],
    'target': [0, 1, 0, 1, 0]
})

# Count NaNs per column
print(df.isnull().sum())
# age         2
# income      1
# category    1
# target      0

# Percentage missing
print(df.isnull().mean().round(3) * 100)
# age         40.0
# income      20.0
# category    20.0
# target       0.0

# Which rows have any NaN
print(df[df.isnull().any(axis=1)])
#     age   income category  target
# 1   NaN  60000.0        B       1
# 2  35.0      NaN        A       0
# 3  40.0  80000.0      NaN       1
# 4   NaN  70000.0        B       0

Four NaNs, but they sit in four different rows, so df.dropna() leaves 1 of 5 rows. If you already have a NumPy array instead of a DataFrame:

import numpy as np

X = np.array([[1, 2], [np.nan, 3], [4, np.nan]])

# Which columns have NaN
print(np.isnan(X).any(axis=0))  # [ True  True]

# Count per column
print(np.isnan(X).sum(axis=0))  # [1 1]

# Which rows
print(np.where(np.isnan(X).any(axis=1))[0])  # [1 2]

For the infinity variant, check with np.isinf(df.select_dtypes('number')).sum(). The usual source is a ratio column where the denominator was 0: pd.Series([1., 2, 3, 4]) / pd.Series([1., 0, 2, 1]) gives [1.0, inf, 1.5, 4.0].


Step 2: Choose an Imputation Strategy

sklearn's SimpleImputer covers four strategies:

  • mean: replace NaN with the column mean. Sensitive to outliers.
  • median: replace NaN with the column median. Safer when the feature is skewed or has outliers.
  • most_frequent: replace NaN with the most common value. Works for categorical columns stored as strings.
  • constant: replace NaN with a value you pass via fill_value, e.g. 0 for "no purchases" or "Unknown" for a missing category.
from sklearn.impute import SimpleImputer
import numpy as np

X = np.array([[1, 2], [np.nan, 3], [4, np.nan], [6, 8]])

print(SimpleImputer(strategy='mean').fit_transform(X))
# [[1.         2.        ]
#  [3.66666667 3.        ]    NaN -> (1+4+6)/3
#  [4.         4.33333333]    NaN -> (2+3+8)/3
#  [6.         8.        ]]

print(SimpleImputer(strategy='median').fit_transform(X))
# [[1. 2.]
#  [4. 3.]
#  [4. 3.]
#  [6. 8.]]

print(SimpleImputer(strategy='constant', fill_value=0).fit_transform(X))
# [[1. 2.]
#  [0. 3.]
#  [4. 0.]
#  [6. 8.]]

If the fact that a value was missing might itself be predictive, add add_indicator=True. It appends one 0/1 column per feature that had NaN during fit:

print(SimpleImputer(strategy='mean', add_indicator=True).fit_transform(X))
# [[1.         2.         0.         0.        ]
#  [3.66666667 3.         1.         0.        ]
#  [4.         4.33333333 0.         1.        ]
#  [6.         8.         0.         0.        ]]

KNNImputer and IterativeImputer

These use the other columns to estimate the missing value instead of a single column statistic. On the same X:

from sklearn.impute import KNNImputer
print(KNNImputer(n_neighbors=2).fit_transform(X))
# [[1.  2. ]
#  [3.5 3. ]
#  [4.  5. ]
#  [6.  8. ]]

# IterativeImputer is still experimental in 1.9.1; importing it directly raises ImportError
from sklearn.experimental import enable_iterative_imputer  # noqa: F401
from sklearn.impute import IterativeImputer
print(IterativeImputer(random_state=0).fit_transform(X).round(3))
# [[1.    2.   ]
#  [1.834 3.   ]
#  [4.    5.6  ]
#  [6.    8.   ]]

Both drop into a Pipeline exactly like SimpleImputer. KNNImputer uses distances, so it is affected by feature scale; on a toy array that's fine, on real data with income next to age it isn't.


Step 3: Use SimpleImputer Inside a Pipeline, Not Before

The most common mistake is fitting the imputer on the full dataset before the train/test split. The test rows then contribute to the mean or median used to fill training rows, so your evaluation is no longer on truly unseen data.

The Wrong Way

from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
import numpy as np

rng = np.random.RandomState(0)
X = rng.randn(1000, 10)
X[rng.rand(*X.shape) < 0.1] = np.nan
y = rng.randint(0, 2, 1000)

# BAD: imputer sees test data before split
imputer = SimpleImputer(strategy='mean')
X_filled = imputer.fit_transform(X)  # uses entire X including future test rows

X_train, X_test, y_train, y_test = train_test_split(X_filled, y, test_size=0.2)

model = LogisticRegression()
model.fit(X_train, y_train)
# The fill values were computed with the test rows included

The Correct Way: Pipeline

from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split
import numpy as np

rng = np.random.RandomState(0)
X = rng.randn(1000, 10)
X[rng.rand(*X.shape) < 0.1] = np.nan
y = rng.randint(0, 2, 1000)

# Split FIRST, before any fitting
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Pipeline: imputer fitted only on training data
pipeline = Pipeline([
    ('imputer', SimpleImputer(strategy='median')),
    ('scaler', StandardScaler()),
    ('model', LogisticRegression(max_iter=500))
])

pipeline.fit(X_train, y_train)

# When predicting, the pipeline applies the same imputation
# (fitted on train stats only) to the test set
score = pipeline.score(X_test, y_test)
print(f"Test accuracy: {score:.3f}")
# Test accuracy: 0.520   (labels are random here, so ~0.5 is expected)

Inside a Pipeline, fit() on training data calls fit_transform() on each step. predict() or score() on test data calls only transform() using the statistics learned from training. The test set never influences the imputer's mean or median. Also note that StandardScaler does not raise on NaN, it passes it through, so a pipeline with no imputer only fails at the model step, which can make the error look like a model problem.


make_pipeline Shorthand

For quick pipelines where you do not need to name the steps, make_pipeline auto-names them:

from sklearn.pipeline import make_pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

# Equivalent to the Pipeline above, steps named automatically
pipeline = make_pipeline(
    SimpleImputer(strategy='median'),
    StandardScaler(),
    LogisticRegression(max_iter=500)
)

pipeline.fit(X_train, y_train)
print(pipeline.score(X_test, y_test))

# Step names are lowercased class names
print(pipeline.named_steps.keys())
# dict_keys(['simpleimputer', 'standardscaler', 'logisticregression'])

The named steps are useful for checking what the imputer learned:

imputer_step = pipeline.named_steps['simpleimputer']
print(imputer_step.statistics_.round(3))  # one median per column, from X_train only
# [-0.037  0.007 -0.057 -0.04  -0.016  0.004 -0.003 -0.021 -0.047  0.04 ]

Handling Mixed Types: Numeric and Categorical Together

When your DataFrame has both numeric and categorical columns with NaN, use ColumnTransformer to apply a different imputation strategy to each group:

import pandas as pd
import numpy as np
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

df = pd.DataFrame({
    'age': [25, np.nan, 35, 40, np.nan, 30],
    'income': [50000, 60000, np.nan, 80000, 70000, 55000],
    'category': ['A', 'B', 'A', np.nan, 'B', 'A'],
    'target': [0, 1, 0, 1, 0, 1]
})

X = df.drop('target', axis=1)
y = df['target']

numeric_cols = ['age', 'income']
categorical_cols = ['category']

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)

preprocessor = ColumnTransformer(transformers=[
    ('num', Pipeline([
        ('imputer', SimpleImputer(strategy='median')),
        ('scaler', StandardScaler())
    ]), numeric_cols),
    ('cat', Pipeline([
        ('imputer', SimpleImputer(strategy='most_frequent')),
        ('encoder', OneHotEncoder(handle_unknown='ignore'))
    ]), categorical_cols)
])

full_pipeline = Pipeline([
    ('preprocessor', preprocessor),
    ('model', LogisticRegression(max_iter=200))
])

full_pipeline.fit(X_train, y_train)

# A new row where every feature is missing still predicts without error
new_row = pd.DataFrame({'age': [np.nan], 'income': [np.nan], 'category': [np.nan]})
print(full_pipeline.predict(new_row))  # [1]

Six rows are only enough to show that it runs; the test split here has 2 rows, so don't read anything into its score.


The "infinity or a value too large" Variant

ValueError: Input X contains infinity or a value too large for dtype('float64') comes from the same input check. Imputers won't save you here, because SimpleImputer raised the exact same error when fed np.inf. Convert inf to NaN first, then let the pipeline impute:

df = pd.DataFrame({'r': pd.Series([1., 2, 3, 4]) / pd.Series([1., 0, 2, 1]),
                   'b': [1., np.nan, 3, 4]})
# df['r'] is [1.0, inf, 1.5, 4.0]

df = df.replace([np.inf, -np.inf], np.nan)
pipe = make_pipeline(SimpleImputer(strategy='median'), LogisticRegression())
pipe.fit(df, [0, 1, 0, 1])   # fits without error

If the message says dtype('float32') and there is no inf in your data, look for huge finite values. RandomForestClassifier converts X to float32, so a value of 1e39 (above float32's max of about 3.4e38) raised it, while LogisticRegression fitted fine on 1e300 in float64.


NaN in the Target: "Input y contains NaN"

ValueError: Input y contains NaN. is a one-line message with no hint about imputers. Every estimator tested, including HistGradientBoostingClassifier and RandomForestClassifier, rejected it. Filling a label with the mean invents training targets. Drop those rows before splitting:

df = df.dropna(subset=['target'])   # only rows with a missing label

Estimators That Accept NaN Directly

The error message only names HistGradientBoostingClassifier and HistGradientBoostingRegressor, but tree support has grown. Result of fit() on a small array with NaN in both columns:

Estimator1.3.21.5.21.9.1
DecisionTreeClassifierfitsfitsfits
RandomForestClassifierInput X contains NaNfitsfits
ExtraTreesClassifierInput X contains NaNInput X contains NaNfits

In 1.9.1 the regressor versions (DecisionTreeRegressor, RandomForestRegressor), ExtraTreeClassifier and HistGradientBoostingClassifier also fitted on NaN, and every tree model listed predicted on a row that was entirely NaN. GradientBoostingClassifier (the non-Hist one), linear models, SVC and KNeighborsClassifier did not. If you're on an older release, upgrading may be the whole fix; check with python -c "import sklearn; print(sklearn.__version__)". HistGradientBoostingClassifier also fitted with np.inf in X, while RandomForestClassifier raised the float32 error.


When to Drop vs Impute

Dropping rows (df.dropna()) looks harmless until you count. With 10 columns and 10% of values missing at random, np.isnan(X).any(axis=1).sum() reported 683 of 1000 rows with at least one NaN: dropna() would throw away 68% of the data to remove 10% of the values.

Dropping is reasonable when:

  • Only a small fraction of rows are affected and you have data to spare.
  • The missing value is the target (see above).

Otherwise impute, and with add_indicator=True if missingness might carry signal. For a broader overview of how preprocessing choices interact with model choice, see the top 10 ML algorithms guide.

fillna() pitfalls

  • df.fillna(df.mean()) on the whole DataFrame is the same leakage as the wrong-way example above: the mean includes test rows. On [1, NaN, 3, 100] it filled 34.67, pulled up by the single outlier; SimpleImputer(strategy='median') in a pipeline avoids both problems.
  • df.fillna(0) on a string column produced ['A', 0, 'B']: an int mixed into strings. OneHotEncoder then raised TypeError: Encoders require their input argument must be uniformly strings or numbers. Got ['int', 'str']. Use a string such as 'Unknown' for text columns.
  • df.fillna(method='ffill') from older tutorials raised TypeError: NDFrame.fillna() got an unexpected keyword argument 'method' on pandas 3.0.6. Use df.ffill() or df.bfill().

If you see a ConvergenceWarning after imputing, your features probably have very different scales; see the ConvergenceWarning fix guide and keep a scaler in the Pipeline after the imputer.