Use str.replace() to change matching text within a Pandas column. Choose literal or regex matching explicitly, preserve missing values and check the result before overwriting source data. Use DataFrame.replace() for whole-value and numeric replacements.
Keep the source data available so each replacement rule can be compared independently.
import pandas as pd
my_dict={
'id':[1,2,3,4,5,4,2],
'name':['John','Max','Arnold','Krish','John','Krish','Max'],
'class1':['Four','Three','Three','Four','Four','Four','Three'],
'mark':[75,85,55,60,60,60,85],
'gender':['female','male','male','female','female','female','male']
}
df = pd.DataFrame(data=my_dict)
print(df)Expected output
id name class1 mark gender
0 1 John Four 75 female
1 2 Max Three 85 male
2 3 Arnold Three 55 male
3 4 Krish Four 60 female
4 5 John Four 60 female
5 4 Krish Four 60 female
6 2 Max Three 85 malestr.replace() edits matching text inside each string in one Series. DataFrame.replace() replaces whole matching values by default and can also apply regex to string values when requested. Numeric replacement belongs to replace(), not the string accessor. Assign the returned result to keep the change. In current Pandas, str.replace() defaults to regex=False; explicitly choose True for patterns.
Without regex, DataFrame.replace() replaces an entire matching cell, rather than a substring.
df = pd.DataFrame(my_dict)
result = df.replace('Max', 'Jim')
print(result)Expected output
id name class1 mark gender
0 1 John Four 75 female
1 2 Jim Three 85 male
2 3 Arnold Three 55 male
3 4 Krish Four 60 female
4 5 John Four 60 female
5 4 Krish Four 60 female
6 2 Jim Three 85 maleThis table-wide replacement changes 85 wherever that value occurs. For a business-specific correction, target the intended column.
df = pd.DataFrame(my_dict)
print(df.replace(85, 100))Expected output
id name class1 mark gender
0 1 John Four 75 female
1 2 Max Three 100 male
2 3 Arnold Three 55 male
3 4 Krish Four 60 female
4 5 John Four 60 female
5 4 Krish Four 60 female
6 2 Max Three 100 maleThe old and new lists are paired in order and must have matching lengths.
df = pd.DataFrame(my_dict)
print(df.replace(['John', 85], ['Jim', 100]))Expected output
id name class1 mark gender
0 1 Jim Four 75 female
1 2 Max Three 100 male
2 3 Arnold Three 55 male
3 4 Krish Four 60 female
4 5 Jim Four 60 female
5 4 Krish Four 60 female
6 2 Max Three 100 maleA flat dictionary maps old values to new values throughout the table.
df = pd.DataFrame(my_dict)
print(df.replace({'John': 'Jim', 85: 100}))Expected output
id name class1 mark gender
0 1 Jim Four 75 female
1 2 Max Three 100 male
2 3 Arnold Three 55 male
3 4 Krish Four 60 female
4 5 Jim Four 60 female
5 4 Krish Four 60 female
6 2 Max Three 100 maleThis demonstrates matching several existing values; replacing marks with 100 is a practice rule, not a recommended correction policy.
df = pd.DataFrame(my_dict)
print(df.replace([75, 85, 60], 100))Expected output
id name class1 mark gender
0 1 John Four 100 female
1 2 Max Three 100 male
2 3 Arnold Three 55 male
3 4 Krish Four 100 female
4 5 John Four 100 female
5 4 Krish Four 100 female
6 2 Max Three 100 maleHere John in name and Four in class1 become Jim. The syntax is useful for column-specific rules, although real replacements should preserve the meaning of each field.
df = pd.DataFrame(my_dict)
print(df.replace({'name': 'John', 'class1': 'Four'}, 'Jim'))Expected output
id name class1 mark gender
0 1 Jim Jim 75 female
1 2 Max Three 85 male
2 3 Arnold Three 55 male
3 4 Krish Jim 60 female
4 5 Jim Jim 60 female
5 4 Krish Jim 60 female
6 2 Max Three 85 maleThe string accessor replaces matching occurrences within each value. Literal mode avoids interpreting the search term as a pattern.
df = pd.DataFrame(my_dict)
df['class1'] = df['class1'].str.replace('Three', 'Ten', regex=False)
print(df)Expected output
id name class1 mark gender
0 1 John Four 75 female
1 2 Max Ten 85 male
2 3 Arnold Ten 55 male
3 4 Krish Four 60 female
4 5 John Four 60 female
5 4 Krish Four 60 female
6 2 Max Ten 85 maleThe anchor restricts replacement to the beginning of a string. Numeric columns are not converted to text by this operation.
df = pd.DataFrame(my_dict)
print(df.replace(regex='^[AF]', value='*'))Expected output
id name class1 mark gender
0 1 John *our 75 female
1 2 Max Three 85 male
2 3 *rnold Three 55 male
3 4 Krish *our 60 female
4 5 John *our 60 female
5 4 Krish *our 60 female
6 2 Max Three 85 maleThe pattern starts with M, followed by two characters, and ends there. It matches Max.
df = pd.DataFrame(my_dict)
print(df.replace(regex={r'^M..$': 'foo'}))Expected output
id name class1 mark gender
0 1 John Four 75 female
1 2 foo Three 85 male
2 3 Arnold Three 55 male
3 4 Krish Four 60 female
4 5 John Four 60 female
5 4 Krish Four 60 female
6 2 foo Three 85 maleThe hn suffix in John becomes foo. The rest of the name is retained.
df = pd.DataFrame(my_dict)
print(df.replace(regex={r'hn$': 'foo'}))Expected output
id name class1 mark gender
0 1 Jofoo Four 75 female
1 2 Max Three 85 male
2 3 Arnold Three 55 male
3 4 Krish Four 60 female
4 5 Jofoo Four 60 female
5 4 Krish Four 60 female
6 2 Max Three 85 maleUse literal mode when @ is the exact character to replace. This exercise creates display text, not a valid email-address transformation.
emails = pd.Series(['Ravi@example.com', 'Raju@example.com', 'Alex@example.com'], dtype='string')
print(emails.str.replace('@', '#', regex=False))Expected output
0 Ravi#example.com
1 Raju#example.com
2 Alex#example.com
dtype: stringcase=False matches Ravi even though the search text is lowercase.
print(emails.str.replace('ravi', 'Ronn', case=False, regex=False))Expected output
0 Ronn@example.com
1 Raju@example.com
2 Alex@example.com
dtype: stringRegex mode must be explicit for the pattern to match instead of being treated literally.
emails = pd.Series(['Ra2vi@example.com', 'Raju@example.com', 'Alex@example.com'], dtype='string')
print(emails.str.replace(r'^[AC]', '*', regex=True))Expected output
0 Ra2vi@example.com
1 Raju@example.com
2 *lex@example.com
dtype: stringThe character range matches digits. It changes the 2 in Ra2vi without affecting the other letters.
print(emails.str.replace(r'[0-9]', '*', regex=True))Expected output
0 Ra*vi@example.com
1 Raju@example.com
2 Alex@example.com
dtype: stringUse [ab] for either character. Inside a character class, | is a literal pipe, so [a|b] would also match pipes.
print(emails.str.replace(r'[ab]', '*', regex=True))Expected output
0 R*2vi@ex*mple.com
1 R*ju@ex*mple.com
2 Alex@ex*mple.com
dtype: stringn=1 changes only the first matching occurrence in each string. The default n=-1 replaces all matches.
print(emails.str.replace(r'[ab]', '*', n=1, regex=True))Expected output
0 R*2vi@example.com
1 R*ju@example.com
2 Alex@ex*mple.com
dtype: stringA dot is special in regex mode. Use regex=False for a single literal search term. Replacement applies within a value, so A.B changes without changing AxB.
labels = pd.Series(["A.B", "AxB", "C++ guide", None], dtype="string")
print(labels.str.replace("A.B", "AB", regex=False))
print(labels.str.replace("C++", "C plus plus", regex=False))
assert labels.str.replace("A.B", "AB", regex=False).iloc[1] == "AxB"Expected output
0 AB
1 AxB
2 C++ guide
3 <NA>
dtype: string
0 A.B
1 AxB
2 C plus plus guide
3 <NA>
dtype: stringstr.replace() preserves missing values; it has no na parameter. Convert mixed values to nullable string only when numbers should become searchable text. Otherwise validate or exclude unexpected numeric inputs. See astype() and fillna().
raw = pd.Series([" notebook ", None, 120], dtype="object")
text = raw.astype("string")
clean = text.str.strip().str.replace("notebook", "Notebook", regex=False)
print(clean)
assert clean.isna().sum() == 1
assert clean.iloc[2] == "120"Expected output
0 Notebook
1 <NA>
2 120
dtype: stringNamed capture groups let you rearrange text while keeping the captured values. Use raw strings for both pattern and replacement. A callable replacement can generate a result from each regex match; callable replacements require regex=True.
codes = pd.Series(["A-001", "B-023", None], dtype="string")
print(codes.str.replace(r"(?P<letter>[A-Z])-(?P<number>\d+)", r"\g<number>/\g<letter>", regex=True))
print(codes.str.replace(r"\d+", lambda match: str(int(match.group())), regex=True))Expected output
0 001/A
1 023/B
2 <NA>
dtype: string
0 A-1
1 B-23
2 <NA>
dtype: stringThis synthetic sample uses dot-decimal amounts and commas only as thousands separators. Removing commas would be wrong for decimal-comma input. Preserve raw values and define the locale before cleaning. Normalize spacing in product text without changing identifiers.
raw_sales = pd.DataFrame({
"order_id": [101, 102, 103, 104],
"product": [" Notebook A5 ", "Pen blue", None, "Notebook A4"],
"amount_text": ["$1,250.50", "$3.00", "unknown", "$20.00"]
})
sales = raw_sales.copy()
sales["product"] = sales["product"].astype("string").str.strip().str.replace(r"\s+", " ", regex=True)
amount_text = sales["amount_text"].astype("string").str.replace("$", "", regex=False).str.replace(",", "", regex=False).str.strip()
sales["amount"] = pd.to_numeric(amount_text, errors="coerce")
print(sales[["order_id", "product", "amount"]])
assert sales.loc[0, "amount"] == 1250.5
assert sales.loc[0, "product"] == "Notebook A5"Expected output
order_id product amount
0 101 Notebook A5 1250.5
1 102 Pen blue 3.0
2 103 <NA> <NA>
3 104 Notebook A4 20.0Compare nullable text after filling missing values with the same placeholder only for the comparison. Report invalid inputs before exporting. The known subtotal excludes unknown amounts and must not be presented as a complete total. See sum(), str.contains() and Excel export.
before = raw_sales["product"].astype("string")
changed = before.fillna("").ne(sales["product"].fillna(""))
failed = raw_sales["amount_text"].notna() & sales["amount"].isna()
print("Changed products:", sales.loc[changed, "order_id"].tolist())
print("Invalid amounts:", raw_sales.loc[failed, "amount_text"].tolist())
print("Known subtotal:", sales["amount"].sum(min_count=1))
assert failed.sum() == 1
# Optional export in Colab:
# sales.to_csv("cleaned_sales.csv", index=False)Expected output
Changed products: [101, 102]
Invalid amounts: ['unknown']
Known subtotal: 1273.5Why is my regex unchanged? Pass regex=True. Why did all dots disappear? A regex dot matches almost any character; use literal mode for a dot. Why did the original column stay unchanged? Assign the returned Series. How do I replace a whole cell? Use Series.replace() or DataFrame.replace() without regex. How do I count replacements? str.replace() returns updated text, not counts; count matches separately before changing values. Can I pass flags with a compiled regex? Include flags in the compiled pattern rather than also passing flags or case. Avoid applying untrusted complicated patterns to large inputs.
Create labels " Pen blue ", "Notebook A5" and a missing value. Strip surrounding spaces, collapse internal whitespace and replace the whole word Pen with Pencil, case-insensitively. Keep the missing value missing. Expect Pencil blue and Notebook A5.
The word boundary limits replacement to the label Pen, rather than matching part of another word.
labels = pd.Series([" Pen blue ", "Notebook A5", None], dtype="string")
result = labels.str.strip().str.replace(r"\s+", " ", regex=True).str.replace(r"\bPen\b", "Pencil", case=False, regex=True)
print(result)
assert result.iloc[0] == "Pencil blue"
assert result.iloc[1] == "Notebook A5"
assert pd.isna(result.iloc[2])Expected output
0 Pencil blue
1 Notebook A5
2 <NA>
dtype: stringOpen in Google Colab View on GitHub
Run the examples in order, change the replacement rules and compare raw and cleaned values. Save your own copy to keep edits. All sample tables are included in the notebook.
Continue with data cleaning, case conversion, string slicing and splitting text.
Reference: Pandas Series.str.replace documentation.
Author & Instructor at plus2net
I write and maintain practical tutorials on Python, PHP, SQL, JavaScript, HTML, jQuery, and web development at plus2net. The tutorials focus on clear explanations, working examples, and code that readers can test and adapt while learning.