Fix sklearn ValueError: Input X Contains NaN
y) and the estimator can't handle it. Run df.isna().sum() to see which columns, then pick one:
- Impute inside a Pipeline:
make_pipeline(SimpleImputer(strategy='median'), StandardScaler(), LogisticRegression()). Split first, then fit, so test rows never feed the imputer. - Use an estimator that takes NaN: in 1.9.1,
HistGradientBoosting*,DecisionTree*,RandomForest*,ExtraTree(s)*all fit on NaN without complaint. On 1.3.2, ofDecisionTreeClassifier,RandomForestClassifierandExtraTreesClassifier, only the single tree did. - The "infinity" variant:
df.replace([np.inf, -np.inf], np.nan), then impute.SimpleImputeritself rejects inf. - NaN in
y: drop those rows. No estimator accepts it and imputing a target is rarely what you want.
You call model.fit(X, y) and get this (exact text from scikit-learn 1.9.1, the second line is one long line):
ValueError: Input X contains NaN.
LogisticRegression does not accept missing values encoded as NaN natively. For supervised learning, you might want to consider sklearn.ensemble.HistGradientBoostingClassifier and Regressor which accept missing values encoded as NaNs natively. Alternatively, it is possible to preprocess the data, for instance by using an imputer transformer in a pipeline or drop samples with missing values. See https://scikit-learn.org/stable/modules/impute.html You can find a list of all estimators that handle NaN values at the following page: https://scikit-learn.org/stable/modules/impute.html#estimators-that-handle-nan-values
The estimator name changes (LinearRegression, SVC, KNeighborsClassifier, GradientBoostingClassifier all produced the same text with their own name), and so does the call: predict() on a row with NaN raises the identical message even if training data was clean. Dropping every row with a NaN is usually the worst fix. Below: how to find the NaNs, which imputer to use, and how to wire it so test data doesn't leak into training.
Reproduced: Each Error and the Fix That Worked
Every row was run on the versions in the header, error text copied from the output.
| What was passed | Exact error / result | Fix that worked |
|---|---|---|
np.nan in X, LogisticRegression().fit | ValueError: Input X contains NaN. (plus the long hint above) | SimpleImputer in a Pipeline |
Nullable Int64 column with pd.NA | Same Input X contains NaN. message | Same |
np.inf or -np.inf in X (e.g. a ratio divided by 0) | ValueError: Input X contains infinity or a value too large for dtype('float64'). | replace([np.inf, -np.inf], np.nan), then impute |
Finite 1e39 in X, RandomForestClassifier | ValueError: Input X contains infinity or a value too large for dtype('float32'). | Trees cast to float32 (max about 3.4e38): clip or rescale the column |
np.inf into SimpleImputer | ValueError: Input X contains infinity or a value too large for dtype('float64'). | Replace inf before the imputer |
NaN in y (LogisticRegression, LinearRegression, RandomForest, HistGradientBoosting) | ValueError: Input y contains NaN. | Drop rows where y is NaN |
NaN in X, HistGradientBoostingClassifier, RandomForest*, DecisionTree*, ExtraTree(s)Classifier | Fit and predict succeeded (1.9.1) | None needed |
StandardScaler().fit_transform on NaN | No error: NaN passed straight through to the output | Put the imputer before the scaler |
from sklearn.impute import IterativeImputer | ImportError: IterativeImputer is experimental and the API might change without any deprecation cycle. To use it, you need to explicitly import enable_iterative_imputer: | from sklearn.experimental import enable_iterative_imputer first |
Old tutorial code: df.fillna(method='ffill') on pandas 3.0.6 | TypeError: NDFrame.fillna() got an unexpected keyword argument 'method' | df.ffill() |
Note that a very large but finite value like 1e300 did not raise with LogisticRegression: the "too large" part only bites when the estimator converts to a smaller dtype, as tree models do with float32.
Step 1: Find Which Columns Have NaN
The error tells you NaN exists but not where. Locate it first.
import pandas as pd
import numpy as np
df = pd.DataFrame({
'age': [25, np.nan, 35, 40, np.nan],
'income': [50000, 60000, np.nan, 80000, 70000],
'category': ['A', 'B', 'A', np.nan, 'B'],
'target': [0, 1, 0, 1, 0]
})
# Count NaNs per column
print(df.isnull().sum())
# age 2
# income 1
# category 1
# target 0
# Percentage missing
print(df.isnull().mean().round(3) * 100)
# age 40.0
# income 20.0
# category 20.0
# target 0.0
# Which rows have any NaN
print(df[df.isnull().any(axis=1)])
# age income category target
# 1 NaN 60000.0 B 1
# 2 35.0 NaN A 0
# 3 40.0 80000.0 NaN 1
# 4 NaN 70000.0 B 0
Four NaNs, but they sit in four different rows, so df.dropna() leaves 1 of 5 rows. If you already have a NumPy array instead of a DataFrame:
import numpy as np
X = np.array([[1, 2], [np.nan, 3], [4, np.nan]])
# Which columns have NaN
print(np.isnan(X).any(axis=0)) # [ True True]
# Count per column
print(np.isnan(X).sum(axis=0)) # [1 1]
# Which rows
print(np.where(np.isnan(X).any(axis=1))[0]) # [1 2]
For the infinity variant, check with np.isinf(df.select_dtypes('number')).sum(). The usual source is a ratio column where the denominator was 0: pd.Series([1., 2, 3, 4]) / pd.Series([1., 0, 2, 1]) gives [1.0, inf, 1.5, 4.0].
Step 2: Choose an Imputation Strategy
sklearn's SimpleImputer covers four strategies:
- mean: replace NaN with the column mean. Sensitive to outliers.
- median: replace NaN with the column median. Safer when the feature is skewed or has outliers.
- most_frequent: replace NaN with the most common value. Works for categorical columns stored as strings.
- constant: replace NaN with a value you pass via
fill_value, e.g. 0 for "no purchases" or "Unknown" for a missing category.
from sklearn.impute import SimpleImputer
import numpy as np
X = np.array([[1, 2], [np.nan, 3], [4, np.nan], [6, 8]])
print(SimpleImputer(strategy='mean').fit_transform(X))
# [[1. 2. ]
# [3.66666667 3. ] NaN -> (1+4+6)/3
# [4. 4.33333333] NaN -> (2+3+8)/3
# [6. 8. ]]
print(SimpleImputer(strategy='median').fit_transform(X))
# [[1. 2.]
# [4. 3.]
# [4. 3.]
# [6. 8.]]
print(SimpleImputer(strategy='constant', fill_value=0).fit_transform(X))
# [[1. 2.]
# [0. 3.]
# [4. 0.]
# [6. 8.]]
If the fact that a value was missing might itself be predictive, add add_indicator=True. It appends one 0/1 column per feature that had NaN during fit:
print(SimpleImputer(strategy='mean', add_indicator=True).fit_transform(X))
# [[1. 2. 0. 0. ]
# [3.66666667 3. 1. 0. ]
# [4. 4.33333333 0. 1. ]
# [6. 8. 0. 0. ]]
KNNImputer and IterativeImputer
These use the other columns to estimate the missing value instead of a single column statistic. On the same X:
from sklearn.impute import KNNImputer
print(KNNImputer(n_neighbors=2).fit_transform(X))
# [[1. 2. ]
# [3.5 3. ]
# [4. 5. ]
# [6. 8. ]]
# IterativeImputer is still experimental in 1.9.1; importing it directly raises ImportError
from sklearn.experimental import enable_iterative_imputer # noqa: F401
from sklearn.impute import IterativeImputer
print(IterativeImputer(random_state=0).fit_transform(X).round(3))
# [[1. 2. ]
# [1.834 3. ]
# [4. 5.6 ]
# [6. 8. ]]
Both drop into a Pipeline exactly like SimpleImputer. KNNImputer uses distances, so it is affected by feature scale; on a toy array that's fine, on real data with income next to age it isn't.
Step 3: Use SimpleImputer Inside a Pipeline, Not Before
The most common mistake is fitting the imputer on the full dataset before the train/test split. The test rows then contribute to the mean or median used to fill training rows, so your evaluation is no longer on truly unseen data.
The Wrong Way
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
import numpy as np
rng = np.random.RandomState(0)
X = rng.randn(1000, 10)
X[rng.rand(*X.shape) < 0.1] = np.nan
y = rng.randint(0, 2, 1000)
# BAD: imputer sees test data before split
imputer = SimpleImputer(strategy='mean')
X_filled = imputer.fit_transform(X) # uses entire X including future test rows
X_train, X_test, y_train, y_test = train_test_split(X_filled, y, test_size=0.2)
model = LogisticRegression()
model.fit(X_train, y_train)
# The fill values were computed with the test rows included
The Correct Way: Pipeline
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split
import numpy as np
rng = np.random.RandomState(0)
X = rng.randn(1000, 10)
X[rng.rand(*X.shape) < 0.1] = np.nan
y = rng.randint(0, 2, 1000)
# Split FIRST, before any fitting
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Pipeline: imputer fitted only on training data
pipeline = Pipeline([
('imputer', SimpleImputer(strategy='median')),
('scaler', StandardScaler()),
('model', LogisticRegression(max_iter=500))
])
pipeline.fit(X_train, y_train)
# When predicting, the pipeline applies the same imputation
# (fitted on train stats only) to the test set
score = pipeline.score(X_test, y_test)
print(f"Test accuracy: {score:.3f}")
# Test accuracy: 0.520 (labels are random here, so ~0.5 is expected)
Inside a Pipeline, fit() on training data calls fit_transform() on each step. predict() or score() on test data calls only transform() using the statistics learned from training. The test set never influences the imputer's mean or median. Also note that StandardScaler does not raise on NaN, it passes it through, so a pipeline with no imputer only fails at the model step, which can make the error look like a model problem.
make_pipeline Shorthand
For quick pipelines where you do not need to name the steps, make_pipeline auto-names them:
from sklearn.pipeline import make_pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
# Equivalent to the Pipeline above, steps named automatically
pipeline = make_pipeline(
SimpleImputer(strategy='median'),
StandardScaler(),
LogisticRegression(max_iter=500)
)
pipeline.fit(X_train, y_train)
print(pipeline.score(X_test, y_test))
# Step names are lowercased class names
print(pipeline.named_steps.keys())
# dict_keys(['simpleimputer', 'standardscaler', 'logisticregression'])
The named steps are useful for checking what the imputer learned:
imputer_step = pipeline.named_steps['simpleimputer']
print(imputer_step.statistics_.round(3)) # one median per column, from X_train only
# [-0.037 0.007 -0.057 -0.04 -0.016 0.004 -0.003 -0.021 -0.047 0.04 ]
Handling Mixed Types: Numeric and Categorical Together
When your DataFrame has both numeric and categorical columns with NaN, use ColumnTransformer to apply a different imputation strategy to each group:
import pandas as pd
import numpy as np
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
df = pd.DataFrame({
'age': [25, np.nan, 35, 40, np.nan, 30],
'income': [50000, 60000, np.nan, 80000, 70000, 55000],
'category': ['A', 'B', 'A', np.nan, 'B', 'A'],
'target': [0, 1, 0, 1, 0, 1]
})
X = df.drop('target', axis=1)
y = df['target']
numeric_cols = ['age', 'income']
categorical_cols = ['category']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)
preprocessor = ColumnTransformer(transformers=[
('num', Pipeline([
('imputer', SimpleImputer(strategy='median')),
('scaler', StandardScaler())
]), numeric_cols),
('cat', Pipeline([
('imputer', SimpleImputer(strategy='most_frequent')),
('encoder', OneHotEncoder(handle_unknown='ignore'))
]), categorical_cols)
])
full_pipeline = Pipeline([
('preprocessor', preprocessor),
('model', LogisticRegression(max_iter=200))
])
full_pipeline.fit(X_train, y_train)
# A new row where every feature is missing still predicts without error
new_row = pd.DataFrame({'age': [np.nan], 'income': [np.nan], 'category': [np.nan]})
print(full_pipeline.predict(new_row)) # [1]
Six rows are only enough to show that it runs; the test split here has 2 rows, so don't read anything into its score.
The "infinity or a value too large" Variant
ValueError: Input X contains infinity or a value too large for dtype('float64') comes from the same input check. Imputers won't save you here, because SimpleImputer raised the exact same error when fed np.inf. Convert inf to NaN first, then let the pipeline impute:
df = pd.DataFrame({'r': pd.Series([1., 2, 3, 4]) / pd.Series([1., 0, 2, 1]),
'b': [1., np.nan, 3, 4]})
# df['r'] is [1.0, inf, 1.5, 4.0]
df = df.replace([np.inf, -np.inf], np.nan)
pipe = make_pipeline(SimpleImputer(strategy='median'), LogisticRegression())
pipe.fit(df, [0, 1, 0, 1]) # fits without error
If the message says dtype('float32') and there is no inf in your data, look for huge finite values. RandomForestClassifier converts X to float32, so a value of 1e39 (above float32's max of about 3.4e38) raised it, while LogisticRegression fitted fine on 1e300 in float64.
NaN in the Target: "Input y contains NaN"
ValueError: Input y contains NaN. is a one-line message with no hint about imputers. Every estimator tested, including HistGradientBoostingClassifier and RandomForestClassifier, rejected it. Filling a label with the mean invents training targets. Drop those rows before splitting:
df = df.dropna(subset=['target']) # only rows with a missing label
Estimators That Accept NaN Directly
The error message only names HistGradientBoostingClassifier and HistGradientBoostingRegressor, but tree support has grown. Result of fit() on a small array with NaN in both columns:
| Estimator | 1.3.2 | 1.5.2 | 1.9.1 |
|---|---|---|---|
DecisionTreeClassifier | fits | fits | fits |
RandomForestClassifier | Input X contains NaN | fits | fits |
ExtraTreesClassifier | Input X contains NaN | Input X contains NaN | fits |
In 1.9.1 the regressor versions (DecisionTreeRegressor, RandomForestRegressor), ExtraTreeClassifier and HistGradientBoostingClassifier also fitted on NaN, and every tree model listed predicted on a row that was entirely NaN. GradientBoostingClassifier (the non-Hist one), linear models, SVC and KNeighborsClassifier did not. If you're on an older release, upgrading may be the whole fix; check with python -c "import sklearn; print(sklearn.__version__)". HistGradientBoostingClassifier also fitted with np.inf in X, while RandomForestClassifier raised the float32 error.
When to Drop vs Impute
Dropping rows (df.dropna()) looks harmless until you count. With 10 columns and 10% of values missing at random, np.isnan(X).any(axis=1).sum() reported 683 of 1000 rows with at least one NaN: dropna() would throw away 68% of the data to remove 10% of the values.
Dropping is reasonable when:
- Only a small fraction of rows are affected and you have data to spare.
- The missing value is the target (see above).
Otherwise impute, and with add_indicator=True if missingness might carry signal. For a broader overview of how preprocessing choices interact with model choice, see the top 10 ML algorithms guide.
fillna() pitfalls
df.fillna(df.mean())on the whole DataFrame is the same leakage as the wrong-way example above: the mean includes test rows. On[1, NaN, 3, 100]it filled 34.67, pulled up by the single outlier;SimpleImputer(strategy='median')in a pipeline avoids both problems.df.fillna(0)on a string column produced['A', 0, 'B']: an int mixed into strings.OneHotEncoderthen raisedTypeError: Encoders require their input argument must be uniformly strings or numbers. Got ['int', 'str']. Use a string such as'Unknown'for text columns.df.fillna(method='ffill')from older tutorials raisedTypeError: NDFrame.fillna() got an unexpected keyword argument 'method'on pandas 3.0.6. Usedf.ffill()ordf.bfill().
If you see a ConvergenceWarning after imputing, your features probably have very different scales; see the ConvergenceWarning fix guide and keep a scaler in the Pipeline after the imputer.
Where to next
Debugging Production ML ยท step 2 of 5