sklearn ConvergenceWarning: What It Means and How to Fix It
- Scale the features inside a Pipeline:
make_pipeline(StandardScaler(), LogisticRegression()). On the unscaled breast cancer dataset, default lbfgs hit its 100-iteration cap; with scaling it converged in 18 iterations (4.7 ms) and scored best: 0.958 test accuracy, 0.981 in 5-fold CV. - Only then raise
max_iter, sized fromn_iter_. On unscaled data lbfgs needed 2607 iterations (about 470 ms, 100x slower) and still scored lower: 0.951 in CV. More iterations remove the warning, not the cause. - Switch solver if you cannot scale:
newton-choleskyconverged in 9 iterations andliblinearin 24 on the same unscaled data. Do not reach forsagaon small data; it needed 525 iterations even after scaling. - Leave
tolalone.tol=0.1made the warning go away and dropped accuracy to 0.923. ChangingCdid not remove the warning at all.
You ran a model, it produced predictions, and buried somewhere in your output is this (exact text from scikit-learn 1.9.1):
ConvergenceWarning: lbfgs failed to converge after 100 iteration(s) (status=1):
STOP: TOTAL NO. OF ITERATIONS REACHED LIMIT
Increase the number of iterations to improve the convergence (max_iter=100).
You might also want to scale the data as shown in:
https://scikit-learn.org/stable/modules/preprocessing.html
Please also refer to the documentation for alternative solver options:
https://scikit-learn.org/stable/modules/linear_model.html#logistic-regression
The exact wording differs between releases, but the meaning is the same. The model did not crash, so the warning is easy to ignore. You should not ignore it: the optimizer stopped before reaching the minimum of the loss, so the coefficients are not the ones the model was supposed to learn. The message lists three fixes. They are not equally good, and the one people usually try first (more iterations) is the weakest.
Reproduced: What Each Fix Actually Did
The test: load_breast_cancer() (569 rows, 30 features, left unscaled; column maxima range from 0.03 to 4254), split 75/25 with stratify=y, random_state=0. Fit time is the median of 5 fits. The test set is only 143 rows, so a 0.007 difference in accuracy is one sample; the 5-fold CV numbers below the table are the more reliable comparison.
| LogisticRegression setup | Warning | n_iter_ | Test acc. | Fit time |
|---|---|---|---|---|
Default (lbfgs, max_iter=100), unscaled | yes | 100 (cap) | 0.930 | 17 to 22 ms |
max_iter=1000, unscaled | yes | 1000 (cap) | 0.944 | 206 ms |
max_iter=5000, unscaled | no | 2607 | 0.937 | 466 to 470 ms |
StandardScaler + default | no | 18 | 0.958 | 4.7 ms |
solver="newton-cholesky", unscaled | no | 9 | 0.937 | 4.3 ms |
solver="liblinear", unscaled | no | 24 | 0.937 | 4.5 ms |
solver="saga", unscaled | yes | 100 (cap) | 0.923 | 16 ms |
solver="saga", scaled, max_iter=10000 | no | 525 | 0.958 | 77 ms |
tol=0.1, unscaled | no | 49 | 0.923 | 10 ms |
C=0.01 or C=100, unscaled | yes | 100 (cap) | 0.923 / 0.937 | 19 ms |
5-fold cross-validation accuracy on the full dataset: unscaled lbfgs scored 0.944 at max_iter=100, 0.953 at 500 and 1000, and 0.951 at 5000. The scaled pipeline with default settings scored 0.981 ± 0.007. Raising max_iter never closed that gap.
Why a converged unscaled model still scores lower: the L2 penalty (controlled by C) treats every coefficient equally. When one feature is in the thousands and another is below 0.1, the same penalty is far stronger on some features than on others. Scaling does not just speed up the optimizer; it changes which model gets fit, usually for the better.
Which Models Emit ConvergenceWarning
Any sklearn estimator that uses an iterative optimizer can emit this warning. The most common ones are:
- LogisticRegression: solvers lbfgs (default), newton-cholesky, newton-cg, liblinear, sag, saga. sag and saga print a shorter message:
The max_iter was reached which means the coef_ did not converge. - LinearSVC / SVC: the liblinear/libsvm optimizer has its own iteration limit. On the breast cancer data,
LinearSVC()converged in 20 iterations even unscaled, so it is not always the culprit. - MLPClassifier / MLPRegressor: with adam or sgd the message is
Stochastic Optimizer: Maximum iterations (200) reached and the optimization hasn't converged yet.Withsolver="lbfgs"it is the lbfgs message above. - SGDClassifier / SGDRegressor: stochastic gradient descent with
max_iter. - Lasso, ElasticNet: coordinate descent. The message is
Objective did not converge. You might want to increase the number of iterations, check the scale of the features or consider increasing regularisation. Duality gap: ...(plainRidgeuses a direct solver by default and does not iterate).
In each case, the optimizer ran until it hit max_iter without meeting its stopping criterion (tol). The warning means the algorithm was cut off mid-optimization.
Root Cause 1: Too Few Iterations
The default max_iter for LogisticRegression is 100. Raising it does make the warning go away, and checking n_iter_ tells you how many iterations the optimizer really used:
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
X, y = load_breast_cancer(return_X_y=True)
model = LogisticRegression() # ConvergenceWarning, n_iter_ == [100]
model.fit(X, y)
model = LogisticRegression(max_iter=5000)
model.fit(X, y)
print(model.n_iter_) # [2338]
On the full dataset n_iter_ came back as 2338; on the 75% training split used in the table it was 2607. Even max_iter=1000 still warned. The same attribute exists on MLPClassifier, SGDClassifier, LinearSVC and Lasso, so use it to size max_iter from a measurement rather than a guess.
The catch: a count like 2607 is a symptom. If the optimizer needs 23 to 26 times its default budget, the problem is usually the shape of the loss surface, not the budget. On this dataset scaling brought the count to 18 and improved accuracy at the same time (see the table). Raise max_iter after scaling, not instead of it.
Many copies of this fix use make_classification(n_samples=10000, n_features=50) as the "before" example. In 1.9.1 that dataset does not warn: lbfgs converged in 8 iterations, because make_classification already produces features on a similar scale. If you want to see the warning yourself, use real unscaled data like load_breast_cancer().
Root Cause 2: Unscaled Features
This is the most common cause, and the warning text points at it. Gradient-based optimizers work best when features are on roughly the same scale. When one feature spans 0 to 0.03 and another spans 0 to 4254, the loss surface is a long, narrow valley, and the optimizer spends its iterations zig-zagging across it.
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=0)
# Scale inside a Pipeline so the scaler only ever sees training data
pipe = make_pipeline(StandardScaler(), LogisticRegression())
pipe.fit(X_train, y_train)
print(pipe[-1].n_iter_) # [18]
print(pipe.score(X_test, y_test)) # 0.958...
Always scale inside a Pipeline, not before the train/test split. Fitting the scaler on the full dataset before splitting leaks information from the test set into training. See the SimpleImputer leakage trap for an identical pattern with imputation.
Use StandardScaler (zero mean, unit variance) for most cases. Use MinMaxScaler when you need bounded output. Use RobustScaler when your data has significant outliers.
Scaling also helped the neural network, with one trap. MLPClassifier() with default adam on the unscaled data stopped at iteration 102 with no warning and scored 0.923: its early-stop check (n_iter_no_change) fired because the loss had stalled, not because it had found a good minimum. Scaled, the same model hit the 200-iteration cap and did warn, but already scored 0.951; with max_iter=1000 it stopped at 267 and scored 0.958. The silent unscaled run was the worst of the three.
Root Cause 3: Wrong Solver for the Problem
Different solvers have different convergence properties. lbfgs is the default and a good general choice once features are scaled. If you cannot scale (for example, the coefficients must stay in original units), a second-order solver copes far better with badly scaled features:
from sklearn.linear_model import LogisticRegression
# Unscaled breast cancer data: converged in 9 iterations, no warning
model = LogisticRegression(solver="newton-cholesky")
model.fit(X_train, y_train)
Solver guide for LogisticRegression in scikit-learn 1.9:
- lbfgs: default. L2 or no penalty. Sensitive to feature scale (100 iterations was not enough unscaled; 18 was enough scaled).
- newton-cholesky: good when
n_samplesis much larger thann_features. 9 iterations unscaled in this test. - liblinear: fast on small data and supports L1, but binary only. With 3 or more classes, 1.9.1 raises
ValueError: The 'liblinear' solver does not support multiclass classification (n_classes >= 3); wrap it inOneVsRestClassifierif you need it. - saga: supports L1 and ElasticNet and scales to very large datasets, but converges slowly on small ones. Here it hit the 100-iteration cap even after scaling and needed 525 iterations; with L1 (
l1_ratio=1.0) it needed 4725. It is not a quick fix for a small dataset. - sag: like saga, L2 only. It needed 315 iterations on the scaled data.
Three API changes to know about, all checked on 1.9.1:
penaltyis deprecated since 1.8 (removal planned for 1.10). Usel1_ratioinstead:l1_ratio=1.0for L1, a value between 0 and 1 for ElasticNet. Passingpenalty="elasticnet"still works but raises aFutureWarning.n_jobshas had no effect since 1.8 and raises aFutureWarning, so drop it fromLogisticRegression(solver="saga", n_jobs=-1)snippets.multi_classis gone:LogisticRegression(multi_class="ovr")raisesTypeError: ... unexpected keyword argument 'multi_class'. UseOneVsRestClassifier(LogisticRegression())for one-vs-rest.
For MLPClassifier on small datasets, solver="lbfgs" can converge in far fewer iterations than adam, but only on scaled data. With hidden_layer_sizes=(100, 50), max_iter=500, unscaled, lbfgs ran all 500 iterations, took 16.8 s and still warned. Scaled, it converged in 16 iterations (0.8 to 1.2 s) and scored 0.965.
from sklearn.neural_network import MLPClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
model = make_pipeline(
StandardScaler(),
MLPClassifier(hidden_layer_sizes=(100, 50), solver="lbfgs", max_iter=500),
)
model.fit(X_train, y_train)
Lasso behaves differently. With a very small penalty (alpha=1e-4) it warned on both unscaled and scaled data at the default 1000 iterations; scaled it needed 1948 iterations and unscaled 12887. With alpha=1e-2 both converged (109 iterations scaled, 774 unscaled). For coordinate descent, scale first and then check whether your alpha is so small that the model is close to unregularized least squares.
When It Is Safe to Suppress the Warning
There are two legitimate reasons to suppress the warning rather than fix the root cause:
- You are doing a quick exploratory analysis and do not care about optimal parameters yet.
- You have already verified that the model's evaluation metrics are acceptable and additional iterations produce negligible improvement.
To verify that more iterations do not help:
import warnings
import numpy as np
from sklearn.exceptions import ConvergenceWarning
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
for n in [100, 500, 1000, 5000]:
with warnings.catch_warnings():
warnings.simplefilter("ignore", ConvergenceWarning)
scores = cross_val_score(LogisticRegression(max_iter=n), X, y, cv=5)
print(f"max_iter={n:5d}: {np.mean(scores):.4f} +/- {np.std(scores):.4f}")
On the unscaled breast cancer data this printed:
max_iter= 100: 0.9438 +/- 0.0142
max_iter= 500: 0.9526 +/- 0.0142
max_iter= 1000: 0.9526 +/- 0.0142
max_iter= 5000: 0.9508 +/- 0.0180
The score plateaus after 500, which is the signal this check is looking for. But a plateau only tells you more iterations will not help. Run the same check on the scaled pipeline before settling: here it scored 0.9807 ± 0.0065, almost 3 points higher. If you still want to silence the warning, do it narrowly:
import warnings
from sklearn.exceptions import ConvergenceWarning
# Suppress only ConvergenceWarning, not all warnings
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=ConvergenceWarning)
model.fit(X, y)
Prefer the context manager form over warnings.filterwarnings('ignore') at module level. The global form silences other useful warnings, including the FutureWarnings above, for the rest of the process.
Quick Fix Checklist
- Add
StandardScalerin a Pipeline before the model. It fixes most cases: 2607 iterations down to 18 in the test above. - If it still warns, check
n_iter_with a largemax_iter, then setmax_itercomfortably above that number. - If you cannot scale, try
solver="newton-cholesky"(orliblinearfor a binary target). Keepsagafor large datasets or L1/ElasticNet. - Increase
tolonly as a last resort. It loosens the stopping criterion, so the optimizer stops earlier; that is not the same as converging (here: no warning, accuracy down to 0.923). Decreasingtoldoes the opposite and can trigger the warning instead of fixing it.
If you are building a full ML pipeline with preprocessing, see also the top 10 ML algorithms guide for solver trade-offs across model families.