Fix AttributeError: 'DataFrame' object has no attribute 'append' (pandas 2.0)
DataFrame.append() and Series.append() were deprecated in pandas 1.4 and removed in 2.0. For one row, use pd.concat([df, pd.DataFrame([new_row])], ignore_index=True). In a loop, collect rows in a list and call pd.DataFrame(rows) once: at 5,000 rows that took 0.0044 s against 1.37 s for concat in the loop on pandas 3.0.6. Do not switch to df._append(): pandas 2.2.3 suggests it in the error, but it is private and pandas 3.0.6 no longer has it.
You upgraded pandas, or deployed to an environment that already has pandas 2.x, and your script immediately dies with an AttributeError. The method DataFrame.append() that your data pipeline relied on simply does not exist anymore. The replacements below were all run on pandas 1.5.3, 2.2.3 and 3.0.6; which one you want depends on whether you add one row or thousands.
import pandas as pd
df = pd.DataFrame({"name": ["Alice", "Bob"], "score": [88, 92]})
new_row = {"name": "Charlie", "score": 95}
result = df.append(new_row, ignore_index=True)
On pandas 3.0.6 (Python 3.12.3) that prints, with the venv path shortened:
Traceback (most recent call last):
File "/tmp/pdappend/pipeline.py", line 6, in <module>
result = df.append(new_row, ignore_index=True)
^^^^^^^^^
File ".../site-packages/pandas/core/generic.py", line 6194, in __getattr__
return object.__getattribute__(self, name)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AttributeError: 'DataFrame' object has no attribute 'append'
pandas 2.2.3 ends the same traceback with a hint: AttributeError: 'DataFrame' object has no attribute 'append'. Did you mean: '_append'? Ignore it. _append is a private method with no stability promise, and on pandas 3.0.6 df._append(...) raises AttributeError: 'DataFrame' object has no attribute '_append'.
On pandas 1.5.3 the same script still works but prints the deprecation warning:
/tmp/pdappend/pipeline.py:6: FutureWarning: The frame.append method is deprecated and will be removed from pandas in a future version. Use pandas.concat instead.
result = df.append(new_row, ignore_index=True)
The deprecation arrived in pandas 1.4.0 (January 2022) and the method was removed in pandas 2.0.0 (April 2023). Any code that ignored that warning now fails with the AttributeError above. Series.append() went the same way: 1.5.3 warns The series.append method is deprecated, and 2.2.3 and 3.0.6 raise AttributeError: 'Series' object has no attribute 'append'.
Reproduced: pandas 1.5.3, 2.2.3 and 3.0.6
Every code block in this post was run in throwaway virtual environments. pandas 1.5.3 does not build on Python 3.12 (pip fell back to a source build that failed), so it ran on Python 3.11.15.
| What ran | Result | Fix that worked |
|---|---|---|
df.append(new_row, ignore_index=True) | 1.5.3: FutureWarning: The frame.append method is deprecated...2.2.3: AttributeError: ... Did you mean: '_append'?3.0.6: AttributeError: 'DataFrame' object has no attribute 'append' | pd.concat([df, pd.DataFrame([new_row])], ignore_index=True) |
pd.Series([1, 2]).append(pd.Series([3])) | 2.2.3 and 3.0.6: AttributeError: 'Series' object has no attribute 'append' | pd.concat([s1, s2]) |
df._append(new_row, ignore_index=True) | Works silently on 1.5.3 and 2.2.3; 3.0.6: AttributeError: 'DataFrame' object has no attribute '_append' | Do not use it; use pd.concat |
pd.concat([df, new_row]) with a raw dict | All three: TypeError: cannot concatenate object of type '<class 'dict'>'; only Series and DataFrame objs are valid | Wrap it: pd.DataFrame([new_row]) |
5,000 rows, concat in a loop starting from pd.DataFrame(columns=["x", "y"]) | Median 1.37 s on 3.0.6, 1.40 s on 2.2.3, 2.12 s on 1.5.3 (df.append loop on 1.5.3: 3.61 s). Both columns end up object dtype | List of dicts, then pd.DataFrame(rows): median 0.0044 s on 3.0.6, int64 columns |
5,000 rows via df.loc[len(df)] = [...] in a loop | Median 2.30 s on 3.0.6, 3.34 s on 1.5.3 | Same list-then-DataFrame pattern |
Why the error occurs
DataFrame.append() was a convenience wrapper that called pd.concat() internally. Every call to .append() created a brand-new DataFrame by copying both the original data and the new rows into fresh memory. This made it easy to write but expensive to call in a loop: every iteration copies everything accumulated so far, so the total work grows quadratically with the number of rows.
The pandas 2.0.0 release notes list DataFrame.append() and Series.append() among the removed deprecations and point to pd.concat() as the replacement. There is no compatibility alias.
Code usually hits it in one of these places:
- Accumulating rows in a loop: iterating over API responses, file chunks, or database cursor results and calling
df = df.append(row)on each iteration. - Merging two DataFrames vertically:
combined = df1.append(df2)used instead of a proper concat. - Adding a single computed row: appending a summary row (totals, averages) to the bottom of a report DataFrame.
Fix 1: Replace with pd.concat()
pd.concat() is the direct replacement for every use of .append(). It accepts a list of DataFrames (or Series) and concatenates them along an axis. The syntax is slightly more verbose but maps one-to-one with the old .append() behaviour.
Appending a dict (single row)
import pandas as pd
df = pd.DataFrame({"name": ["Alice", "Bob"], "score": [88, 92]})
new_row = {"name": "Charlie", "score": 95}
# Before (pandas < 2.0):
# result = df.append(new_row, ignore_index=True)
# After (pandas 2.0+):
result = pd.concat([df, pd.DataFrame([new_row])], ignore_index=True)
print(result)
# name score
# 0 Alice 88
# 1 Bob 92
# 2 Charlie 95
The key difference: you must wrap the dict in a list and pass it to pd.DataFrame() before concatenating. pd.concat() requires all arguments to be DataFrame or Series objects. Passing the dict itself raises TypeError: cannot concatenate object of type '<class 'dict'>'; only Series and DataFrame objs are valid.
Appending one DataFrame to another
import pandas as pd
df1 = pd.DataFrame({"name": ["Alice", "Bob"], "score": [88, 92]})
df2 = pd.DataFrame({"name": ["Charlie", "Diana"], "score": [95, 78]})
# Before:
# combined = df1.append(df2, ignore_index=True)
# After:
combined = pd.concat([df1, df2], ignore_index=True)
print(combined)
# name score
# 0 Alice 88
# 1 Bob 92
# 2 Charlie 95
# 3 Diana 78
Appending multiple DataFrames at once
One advantage of pd.concat() over the old .append(): you can pass a list of any length, making it easy to collapse a collection of DataFrames in a single call.
import pandas as pd
frames = [
pd.DataFrame({"name": ["Alice"], "score": [88]}),
pd.DataFrame({"name": ["Bob"], "score": [92]}),
pd.DataFrame({"name": ["Charlie"], "score": [95]}),
]
# One call, no intermediate copies
combined = pd.concat(frames, ignore_index=True)
print(combined)
# name score
# 0 Alice 88
# 1 Bob 92
# 2 Charlie 95
Preserving the original index
import pandas as pd
df1 = pd.DataFrame({"score": [88, 92]}, index=["alice", "bob"])
df2 = pd.DataFrame({"score": [95, 78]}, index=["charlie", "diana"])
# Keep original index labels (omit ignore_index)
combined = pd.concat([df1, df2])
print(combined)
# score
# alice 88
# bob 92
# charlie 95
# diana 78
# Or reset to a clean integer index
combined_reset = pd.concat([df1, df2], ignore_index=True)
print(combined_reset)
# score
# 0 88
# 1 92
# 2 95
# 3 78
Fix 2: Build a list, then concat once
If you are accumulating rows inside a loop (the most common performance anti-pattern), the correct fix is not to call pd.concat() on each iteration. That would still produce O(n²) allocations. Instead, collect all rows into a plain Python list and call pd.concat() (or pd.DataFrame()) exactly once after the loop finishes.
Pattern: accumulate dicts, build DataFrame once
import pandas as pd
# Before (pandas < 2.0): slow, now also broken:
# result = pd.DataFrame()
# for page in range(1, 11):
# data = fetch_page(page)
# result = result.append(data, ignore_index=True) # AttributeError
# After: correct and fast:
rows = []
for page in range(1, 11):
# Simulate fetching a page of data
page_rows = [
{"id": page * 10 + i, "value": page * i}
for i in range(5)
]
rows.extend(page_rows) # extend the list, not the DataFrame
# Single allocation at the end
result = pd.DataFrame(rows)
print(result.head())
# id value
# 0 10 0
# 1 11 1
# 2 12 2
# 3 13 3
# 4 14 4
Pattern: accumulate DataFrames, concat once
When each iteration produces a chunk DataFrame (e.g., reading CSV chunks or processing batches), collect the chunks and concat at the end:
import pandas as pd
# Simulated chunked processing
def process_chunk(chunk_df: pd.DataFrame) -> pd.DataFrame:
"""Apply some transformation to a chunk."""
chunk_df = chunk_df.copy()
chunk_df["value_squared"] = chunk_df["value"] ** 2
return chunk_df
# Reading a large CSV in chunks
chunks = []
for chunk in pd.read_csv("large_file.csv", chunksize=10_000):
processed = process_chunk(chunk)
chunks.append(processed) # append the chunk DataFrame to a list
# One concat at the end
result = pd.concat(chunks, ignore_index=True)
print(f"Total rows: {len(result):,}")
Pattern: using a generator with concat
pd.concat() also accepts a generator, which saves you from writing the list yourself. It does not save memory: pandas consumes the whole generator into a list before concatenating, so every chunk is still alive at the same time.
import pandas as pd
def generate_dataframes():
for i in range(100):
yield pd.DataFrame({
"batch": [i] * 5,
"value": range(i * 5, i * 5 + 5)
})
# pandas turns the generator into a list, then concatenates once
result = pd.concat(generate_dataframes(), ignore_index=True)
print(result.shape) # (500, 2)
Fix 3: Use loc or iloc for single-row appends
When you genuinely need to add one row to an existing DataFrame in place, for example a totals row on a report, .loc assignment is the most direct approach. You keep the same df object, although pandas still reallocates its columns internally to make room.
Append a single row with loc
import pandas as pd
df = pd.DataFrame({
"region": ["North", "South", "East", "West"],
"revenue": [120_000, 98_000, 145_000, 87_000],
"units": [450, 380, 510, 320],
})
# Compute totals
totals = {
"region": "TOTAL",
"revenue": df["revenue"].sum(),
"units": df["units"].sum(),
}
# Append totals row using loc with the next integer index
next_idx = len(df)
df.loc[next_idx] = totals
print(df)
# region revenue units
# 0 North 120000 450
# 1 South 98000 380
# 2 East 145000 510
# 3 West 87000 320
# 4 TOTAL 450000 1660
Append a row using a named index
import pandas as pd
df = pd.DataFrame(
{"q1": [10, 20, 30], "q2": [15, 25, 35], "q3": [12, 22, 32]},
index=["product_A", "product_B", "product_C"]
)
# Add a row for a new product
df.loc["product_D"] = {"q1": 18, "q2": 28, "q3": 19}
print(df)
# q1 q2 q3
# product_A 10 15 12
# product_B 20 25 22
# product_C 30 35 32
# product_D 18 28 19
On pandas 1.x and 2.x, writing into a DataFrame that is a slice of another one can trigger a SettingWithCopyWarning; call .copy() first: df = original[mask].copy(). pandas 3.0 uses Copy-on-Write and no longer emits that warning. More important, .loc in a loop is not a fast replacement for .append(): adding 5,000 rows one at a time took a median 2.30 s on pandas 3.0.6, slower than concat in a loop. Use it only for a small, fixed number of rows.
Append a row with pd.DataFrame and concat: safe alternative
import pandas as pd
df = pd.DataFrame({
"region": ["North", "South", "East", "West"],
"revenue": [120_000, 98_000, 145_000, 87_000],
})
totals_row = pd.DataFrame([{
"region": "TOTAL",
"revenue": df["revenue"].sum(),
}])
# Clean, no mutation, no SettingWithCopyWarning
df_with_totals = pd.concat([df, totals_row], ignore_index=True)
print(df_with_totals)
# region revenue
# 0 North 120000
# 1 South 98000
# 2 East 145000
# 3 West 87000
# 4 TOTAL 450000
Other pandas 2.0 changes that break old code
DataFrame.append() is a common removal, but pandas 2.0 shipped other breaking changes too. If your codebase used .append(), it was likely written before 2023 and may be hitting several of these at once.
iteritems() removed
DataFrame.iteritems() and Series.iteritems() were removed (2.2.3 and 3.0.6 raise AttributeError: 'DataFrame' object has no attribute 'iteritems'). They were identical to items(), so replace every occurrence with .items():
import pandas as pd
df = pd.DataFrame({"a": [1, 2], "b": [3, 4]})
# Before (raises AttributeError in pandas 2.0):
# for col_name, col_data in df.iteritems():
# print(col_name, col_data.sum())
# After:
for col_name, col_data in df.items():
print(col_name, col_data.sum())
# a 3
# b 7
See also: Fix AttributeError: 'DataFrame' object has no attribute 'iteritems' for a full migration guide.
swaplevel(): not a breaking change
Some migration lists claim that the positional axis argument of DataFrame.swaplevel() was removed. It was not: df.swaplevel(0, 1, 0) ran without an error or warning on pandas 1.5.3, 2.2.3 and 3.0.6. Writing axis=0 as a keyword is still clearer.
Missing values in integer columns: still float64 by default
Integer columns with a missing value still become float64 by default: pd.DataFrame({"count": [1, 2, None, 4]}).dtypes printed float64 on 1.5.3, 2.2.3 and 3.0.6. The nullable Int64 dtype keeps them as integers, but you have to ask for it (or use dtype_backend="numpy_nullable" in the readers):
import pandas as pd
import numpy as np
# Default in 1.5.3, 2.2.3 and 3.0.6: integers with None become float64
# Opt in to the nullable integer dtype explicitly:
df = pd.DataFrame({"count": pd.array([1, 2, None, 4], dtype=pd.Int64Dtype())})
print(df.dtypes)
# count Int64
# dtype: object
print(df["count"].sum()) # 7 (NaN skipped)
# If you need to fill NaN before converting back to plain int:
df["count_filled"] = df["count"].fillna(0).astype(int)
read_csv dtype_backend parameter
pandas 2.0 introduced the dtype_backend parameter to pd.read_csv(), pd.read_parquet(), and related readers. The default remains NumPy-backed dtypes, but you may see code that explicitly sets dtype_backend="numpy_nullable" or "pyarrow". Code that passes it fails on pandas 1.x:
import pandas as pd
# This parameter did not exist before pandas 2.0:
df = pd.read_csv("data.csv", dtype_backend="numpy_nullable")
print(df.dtypes)
# 3.0.6 (data.csv with an int column a and a text column b): a -> Int64, b -> string
# 1.5.3: TypeError: read_csv() got an unexpected keyword argument 'dtype_backend'
Copy-on-Write is the default in pandas 3.0
pandas 2.0 introduced Copy-on-Write (CoW) as an opt-in preview, and in pandas 3.0 it is always on. Code that mutates a column pulled out of a DataFrame and expects the parent to change silently stops working:
import pandas as pd
df = pd.DataFrame({"a": [1, 2, 3], "b": [4, 5, 6]})
col = df["a"]
col.iloc[0] = 99
print(df["a"].tolist())
# 1.5.3 and 2.2.3 (CoW off): [99, 2, 3]
# 3.0.6 (CoW always on): [1, 2, 3]
On pandas 2.x you can opt in with pd.options.mode.copy_on_write = True to find these spots before upgrading. On 3.0.6 setting that option does nothing except emit Pandas4Warning: The 'mode.copy_on_write' option is deprecated. Copy-on-Write can no longer be disabled (it is always enabled with pandas >= 3.0), and setting the option has no impact.
Removed squeeze parameter from read_csv / groupby
The squeeze=True parameter of pd.read_csv() and DataFrame.groupby() was removed. On 2.2.3 and 3.0.6 it raises TypeError: read_csv() got an unexpected keyword argument 'squeeze' (and the same for DataFrame.groupby()). If you used it to automatically return a Series when a single column was selected, replace it with an explicit .squeeze() call or index the column directly:
import pandas as pd
# Before (squeeze= removed in pandas 2.0):
# s = pd.read_csv("single_col.csv", squeeze=True)
# After:
df = pd.read_csv("single_col.csv")
s = df.squeeze() # or df["column_name"] if you know the column name
print(type(s))
# 3.0.6: <class 'pandas.Series'>
# 1.5.3: <class 'pandas.core.series.Series'>
Why list-then-concat is fastest
The quadratic cost of in-loop concatenation
Each call to pd.concat([existing_df, new_row]) inside a loop must:
- Allocate a new block of memory large enough to hold all existing rows plus the new row.
- Copy all existing data into the new block.
- Copy the new row into the new block.
- Release the old block to the garbage collector.
For n rows accumulated one at a time, this copies row 1 a total of n times, row 2 a total of n-1 times, and so on, giving O(n²) total copy operations. A 100,000-row accumulation via in-loop concat performs roughly 5 billion copy operations where only 100,000 are necessary.
import pandas as pd
import time
N = 5_000
# Approach A: in-loop pd.concat (quadratic)
start = time.perf_counter()
df_a = pd.DataFrame(columns=["x", "y"])
for i in range(N):
df_a = pd.concat([df_a, pd.DataFrame([{"x": i, "y": i ** 2}])], ignore_index=True)
time_a = time.perf_counter() - start
# Approach B: list accumulation, single concat (linear)
start = time.perf_counter()
rows = []
for i in range(N):
rows.append({"x": i, "y": i ** 2})
df_b = pd.DataFrame(rows)
time_b = time.perf_counter() - start
print(f"In-loop concat: {time_a:.2f}s")
print(f"List then DataFrame:{time_b:.4f}s")
print(f"Speedup: {time_a / time_b:.0f}x")
# One run of this exact script on pandas 3.0.6 (i5-7500):
# In-loop concat: 1.36s
# List then DataFrame:0.0179s
# Speedup: 76x
#
# Medians of repeated runs (5 loop runs, 21 list runs):
# pandas 3.0.6: 1.37s vs 0.0044s (~310x)
# pandas 2.2.3: 1.40s vs 0.0041s (~340x)
# pandas 1.5.3: 2.12s vs 0.0035s (~600x); df.append loop: 3.61s
The list version is two to three orders of magnitude faster at 5,000 rows, and the gap widens as N grows because the loop is quadratic. The list takes amortised O(1) appends and pandas builds the columns once at the end. Single timings of the fast path are noisy (one run printed 0.0179 s, the median was 0.0044 s), so repeat before trusting any one number.
There is also a dtype trap: starting from pd.DataFrame(columns=["x", "y"]) leaves both columns as object after the loop on all three versions, with no warning. pd.DataFrame(rows) gives int64.
When the size of individual chunks matters
If each iteration produces a chunk DataFrame rather than a single row, collecting them in a Python list and calling pd.concat(chunks) once is still optimal. pd.concat() examines all input shapes upfront and allocates the output once. The gain is smaller than in the row case because each chunk is already large and generating the random data costs the same in both loops.
import pandas as pd
import numpy as np
import time
N_CHUNKS = 200
CHUNK_SIZE = 500
# Approach A: concat inside the loop
start = time.perf_counter()
result_a = pd.DataFrame()
for i in range(N_CHUNKS):
chunk = pd.DataFrame(np.random.randn(CHUNK_SIZE, 4), columns=list("abcd"))
result_a = pd.concat([result_a, chunk], ignore_index=True)
time_a = time.perf_counter() - start
# Approach B: collect, concat once
start = time.perf_counter()
chunks = []
for i in range(N_CHUNKS):
chunk = pd.DataFrame(np.random.randn(CHUNK_SIZE, 4), columns=list("abcd"))
chunks.append(chunk)
result_b = pd.concat(chunks, ignore_index=True)
time_b = time.perf_counter() - start
print(f"In-loop concat: {time_a:.3f}s shape={result_a.shape}")
print(f"Collect + concat:{time_b:.3f}s shape={result_b.shape}")
print(f"Speedup: {time_a / time_b:.1f}x")
# Output on pandas 3.0.6 (i5-7500), two runs:
# In-loop concat: 0.241s shape=(100000, 4)
# Collect + concat:0.043s shape=(100000, 4)
# Speedup: 5.6x
# (second run: 0.237s vs 0.044s, 5.3x; pandas 1.5.3: 5.7x to 8.3x)
Memory usage comparison
The list pattern is not free in memory. A list of Python dicts is much larger than the DataFrame built from it, and both exist at the moment pd.DataFrame(rows) runs. The script below measures it:
import pandas as pd
import tracemalloc
N = 10_000
# Measure peak memory for list-then-DataFrame approach
tracemalloc.start()
rows = [{"x": i, "y": float(i) ** 0.5, "z": i % 7} for i in range(N)]
df = pd.DataFrame(rows)
current, peak = tracemalloc.get_traced_memory()
tracemalloc.stop()
print(f"DataFrame shape: {df.shape}")
print(f"Peak memory: {peak / 1024:.1f} KB")
print(f"DataFrame memory:{df.memory_usage(deep=True).sum() / 1024:.1f} KB")
# Output on pandas 3.0.6:
# DataFrame shape: (10000, 3)
# Peak memory: 3232.6 KB
# DataFrame memory:234.5 KB
Peak traced memory was about 14 times the size of the finished DataFrame (1.5.3 printed 3236.6 KB and the same 234.5 KB). If that matters, collect chunk DataFrames instead of one dict per row, or append to per-column lists and pass a dict of lists to pd.DataFrame.
If you only fix one thing, fix the loop
The AttributeError is the easy part. A find-and-replace from .append() to pd.concat() gets your code running again in five minutes. Before moving on, check whether that code calls append or concat inside a loop. In the runs above that loop cost about 300x at 5,000 rows, and a straight swap to pd.concat keeps the cost and turns the columns into object dtype. Rewriting it to collect a list and build the DataFrame once usually takes a few extra lines.