Pandas str.replace(): Clean Text with Literal and Regex Rules

Use str.replace() to change matching text within a Pandas column. Choose literal or regex matching explicitly, preserve missing values and check the result before overwriting source data. Use DataFrame.replace() for whole-value and numeric replacements.

Create the student dataset 🔝

Keep the source data available so each replacement rule can be compared independently.

import pandas as pd 
my_dict={
  'id':[1,2,3,4,5,4,2],
  'name':['John','Max','Arnold','Krish','John','Krish','Max'],
  'class1':['Four','Three','Three','Four','Four','Four','Three'],
  'mark':[75,85,55,60,60,60,85],
  'gender':['female','male','male','female','female','female','male']
	}
df = pd.DataFrame(data=my_dict)
print(df)

Expected output

   id    name class1  mark  gender
0   1    John   Four    75  female
1   2     Max  Three    85    male
2   3  Arnold  Three    55    male
3   4   Krish   Four    60  female
4   5    John   Four    60  female
5   4   Krish   Four    60  female
6   2     Max  Three    85    male

Choose str.replace() or DataFrame.replace() 🔝

str.replace() edits matching text inside each string in one Series. DataFrame.replace() replaces whole matching values by default and can also apply regex to string values when requested. Numeric replacement belongs to replace(), not the string accessor. Assign the returned result to keep the change. In current Pandas, str.replace() defaults to regex=False; explicitly choose True for patterns.

Replace a whole matching string 🔝

Without regex, DataFrame.replace() replaces an entire matching cell, rather than a substring.

df = pd.DataFrame(my_dict)
result = df.replace('Max', 'Jim')
print(result)

Expected output

   id    name class1  mark  gender
0   1    John   Four    75  female
1   2     Jim  Three    85    male
2   3  Arnold  Three    55    male
3   4   Krish   Four    60  female
4   5    John   Four    60  female
5   4   Krish   Four    60  female
6   2     Jim  Three    85    male

Replace a matching number 🔝

This table-wide replacement changes 85 wherever that value occurs. For a business-specific correction, target the intended column.

df = pd.DataFrame(my_dict)
print(df.replace(85, 100))

Expected output

   id    name class1  mark  gender
0   1    John   Four    75  female
1   2     Max  Three   100    male
2   3  Arnold  Three    55    male
3   4   Krish   Four    60  female
4   5    John   Four    60  female
5   4   Krish   Four    60  female
6   2     Max  Three   100    male

Replace paired lists of values 🔝

The old and new lists are paired in order and must have matching lengths.

df = pd.DataFrame(my_dict)
print(df.replace(['John', 85], ['Jim', 100]))

Expected output

   id    name class1  mark  gender
0   1     Jim   Four    75  female
1   2     Max  Three   100    male
2   3  Arnold  Three    55    male
3   4   Krish   Four    60  female
4   5     Jim   Four    60  female
5   4   Krish   Four    60  female
6   2     Max  Three   100    male

Replace values using a dictionary 🔝

A flat dictionary maps old values to new values throughout the table.

df = pd.DataFrame(my_dict)
print(df.replace({'John': 'Jim', 85: 100}))

Expected output

   id    name class1  mark  gender
0   1     Jim   Four    75  female
1   2     Max  Three   100    male
2   3  Arnold  Three    55    male
3   4   Krish   Four    60  female
4   5     Jim   Four    60  female
5   4   Krish   Four    60  female
6   2     Max  Three   100    male

Replace several numbers with one value 🔝

This demonstrates matching several existing values; replacing marks with 100 is a practice rule, not a recommended correction policy.

df = pd.DataFrame(my_dict)
print(df.replace([75, 85, 60], 100))

Expected output

   id    name class1  mark  gender
0   1    John   Four   100  female
1   2     Max  Three   100    male
2   3  Arnold  Three    55    male
3   4   Krish   Four   100  female
4   5    John   Four   100  female
5   4   Krish   Four   100  female
6   2     Max  Three   100    male

Target different matches in named columns 🔝

Here John in name and Four in class1 become Jim. The syntax is useful for column-specific rules, although real replacements should preserve the meaning of each field.

df = pd.DataFrame(my_dict)
print(df.replace({'name': 'John', 'class1': 'Four'}, 'Jim'))

Expected output

   id    name class1  mark  gender
0   1     Jim    Jim    75  female
1   2     Max  Three    85    male
2   3  Arnold  Three    55    male
3   4   Krish    Jim    60  female
4   5     Jim    Jim    60  female
5   4   Krish    Jim    60  female
6   2     Max  Three    85    male

Replace text in one column 🔝

The string accessor replaces matching occurrences within each value. Literal mode avoids interpreting the search term as a pattern.

df = pd.DataFrame(my_dict)
df['class1'] = df['class1'].str.replace('Three', 'Ten', regex=False)
print(df)

Expected output

   id    name class1  mark  gender
0   1    John   Four    75  female
1   2     Max    Ten    85    male
2   3  Arnold    Ten    55    male
3   4   Krish   Four    60  female
4   5    John   Four    60  female
5   4   Krish   Four    60  female
6   2     Max    Ten    85    male

Apply regex to string values across a table 🔝

The anchor restricts replacement to the beginning of a string. Numeric columns are not converted to text by this operation.

df = pd.DataFrame(my_dict)
print(df.replace(regex='^[AF]', value='*'))

Expected output

   id    name class1  mark  gender
0   1    John   *our    75  female
1   2     Max  Three    85    male
2   3  *rnold  Three    55    male
3   4   Krish   *our    60  female
4   5    John   *our    60  female
5   4   Krish   *our    60  female
6   2     Max  Three    85    male

Match an entire three-character name 🔝

The pattern starts with M, followed by two characters, and ends there. It matches Max.

df = pd.DataFrame(my_dict)
print(df.replace(regex={r'^M..$': 'foo'}))

Expected output

   id    name class1  mark  gender
0   1    John   Four    75  female
1   2     foo  Three    85    male
2   3  Arnold  Three    55    male
3   4   Krish   Four    60  female
4   5    John   Four    60  female
5   4   Krish   Four    60  female
6   2     foo  Three    85    male

Replace matching text at the end 🔝

The hn suffix in John becomes foo. The rest of the name is retained.

df = pd.DataFrame(my_dict)
print(df.replace(regex={r'hn$': 'foo'}))

Expected output

   id    name class1  mark  gender
0   1   Jofoo   Four    75  female
1   2     Max  Three    85    male
2   3  Arnold  Three    55    male
3   4   Krish   Four    60  female
4   5   Jofoo   Four    60  female
5   4   Krish   Four    60  female
6   2     Max  Three    85    male

Replace a literal character in email text 🔝

Use literal mode when @ is the exact character to replace. This exercise creates display text, not a valid email-address transformation.

emails = pd.Series(['Ravi@example.com', 'Raju@example.com', 'Alex@example.com'], dtype='string')
print(emails.str.replace('@', '#', regex=False))

Expected output

0    Ravi#example.com
1    Raju#example.com
2    Alex#example.com
dtype: string

Use case-insensitive replacement 🔝

case=False matches Ravi even though the search text is lowercase.

print(emails.str.replace('ravi', 'Ronn', case=False, regex=False))

Expected output

0    Ronn@example.com
1    Raju@example.com
2    Alex@example.com
dtype: string

Replace an initial A or C 🔝

Regex mode must be explicit for the pattern to match instead of being treated literally.

emails = pd.Series(['Ra2vi@example.com', 'Raju@example.com', 'Alex@example.com'], dtype='string')
print(emails.str.replace(r'^[AC]', '*', regex=True))

Expected output

0    Ra2vi@example.com
1     Raju@example.com
2     *lex@example.com
dtype: string

Replace digits in text 🔝

The character range matches digits. It changes the 2 in Ra2vi without affecting the other letters.

print(emails.str.replace(r'[0-9]', '*', regex=True))

Expected output

0    Ra*vi@example.com
1     Raju@example.com
2     Alex@example.com
dtype: string

Replace a or b characters 🔝

Use [ab] for either character. Inside a character class, | is a literal pipe, so [a|b] would also match pipes.

print(emails.str.replace(r'[ab]', '*', regex=True))

Expected output

0    R*2vi@ex*mple.com
1     R*ju@ex*mple.com
2     Alex@ex*mple.com
dtype: string

Limit replacements per string with n 🔝

n=1 changes only the first matching occurrence in each string. The default n=-1 replaces all matches.

print(emails.str.replace(r'[ab]', '*', n=1, regex=True))

Expected output

0    R*2vi@example.com
1     R*ju@example.com
2     Alex@ex*mple.com
dtype: string

Search for dots and plus signs literally 🔝

A dot is special in regex mode. Use regex=False for a single literal search term. Replacement applies within a value, so A.B changes without changing AxB.

labels = pd.Series(["A.B", "AxB", "C++ guide", None], dtype="string")
print(labels.str.replace("A.B", "AB", regex=False))
print(labels.str.replace("C++", "C plus plus", regex=False))
assert labels.str.replace("A.B", "AB", regex=False).iloc[1] == "AxB"

Expected output

0           AB
1          AxB
2    C++ guide
3         <NA>
dtype: string
0                  A.B
1                  AxB
2    C plus plus guide
3                 <NA>
dtype: string

Preserve missing values and inspect mixed inputs 🔝

str.replace() preserves missing values; it has no na parameter. Convert mixed values to nullable string only when numbers should become searchable text. Otherwise validate or exclude unexpected numeric inputs. See astype() and fillna().

raw = pd.Series([" notebook ", None, 120], dtype="object")
text = raw.astype("string")
clean = text.str.strip().str.replace("notebook", "Notebook", regex=False)
print(clean)
assert clean.isna().sum() == 1
assert clean.iloc[2] == "120"

Expected output

0    Notebook
1        <NA>
2         120
dtype: string

Reuse matched text with capture groups 🔝

Named capture groups let you rearrange text while keeping the captured values. Use raw strings for both pattern and replacement. A callable replacement can generate a result from each regex match; callable replacements require regex=True.

codes = pd.Series(["A-001", "B-023", None], dtype="string")
print(codes.str.replace(r"(?P<letter>[A-Z])-(?P<number>\d+)", r"\g<number>/\g<letter>", regex=True))
print(codes.str.replace(r"\d+", lambda match: str(int(match.group())), regex=True))

Expected output

0    001/A
1    023/B
2     <NA>
dtype: string
0     A-1
1    B-23
2    <NA>
dtype: string

Practical workflow: clean a sales extract 🔝

This synthetic sample uses dot-decimal amounts and commas only as thousands separators. Removing commas would be wrong for decimal-comma input. Preserve raw values and define the locale before cleaning. Normalize spacing in product text without changing identifiers.

raw_sales = pd.DataFrame({
    "order_id": [101, 102, 103, 104],
    "product": ["  Notebook   A5 ", "Pen  blue", None, "Notebook A4"],
    "amount_text": ["$1,250.50", "$3.00", "unknown", "$20.00"]
})
sales = raw_sales.copy()
sales["product"] = sales["product"].astype("string").str.strip().str.replace(r"\s+", " ", regex=True)
amount_text = sales["amount_text"].astype("string").str.replace("$", "", regex=False).str.replace(",", "", regex=False).str.strip()
sales["amount"] = pd.to_numeric(amount_text, errors="coerce")
print(sales[["order_id", "product", "amount"]])
assert sales.loc[0, "amount"] == 1250.5
assert sales.loc[0, "product"] == "Notebook A5"

Expected output

   order_id      product  amount
0       101  Notebook A5  1250.5
1       102     Pen blue     3.0
2       103         <NA>    <NA>
3       104  Notebook A4    20.0

Check what changed and report parsing failures 🔝

Compare nullable text after filling missing values with the same placeholder only for the comparison. Report invalid inputs before exporting. The known subtotal excludes unknown amounts and must not be presented as a complete total. See sum(), str.contains() and Excel export.

before = raw_sales["product"].astype("string")
changed = before.fillna("").ne(sales["product"].fillna(""))
failed = raw_sales["amount_text"].notna() & sales["amount"].isna()
print("Changed products:", sales.loc[changed, "order_id"].tolist())
print("Invalid amounts:", raw_sales.loc[failed, "amount_text"].tolist())
print("Known subtotal:", sales["amount"].sum(min_count=1))
assert failed.sum() == 1
# Optional export in Colab:
# sales.to_csv("cleaned_sales.csv", index=False)

Expected output

Changed products: [101, 102]
Invalid amounts: ['unknown']
Known subtotal: 1273.5

Common questions and mistakes 🔝

Why is my regex unchanged? Pass regex=True. Why did all dots disappear? A regex dot matches almost any character; use literal mode for a dot. Why did the original column stay unchanged? Assign the returned Series. How do I replace a whole cell? Use Series.replace() or DataFrame.replace() without regex. How do I count replacements? str.replace() returns updated text, not counts; count matches separately before changing values. Can I pass flags with a compiled regex? Include flags in the compiled pattern rather than also passing flags or case. Avoid applying untrusted complicated patterns to large inputs.

Exercise: clean product labels 🔝

Create labels " Pen blue ", "Notebook A5" and a missing value. Strip surrounding spaces, collapse internal whitespace and replace the whole word Pen with Pencil, case-insensitively. Keep the missing value missing. Expect Pencil blue and Notebook A5.

Exercise solution 🔝

The word boundary limits replacement to the label Pen, rather than matching part of another word.

labels = pd.Series([" Pen   blue ", "Notebook  A5", None], dtype="string")
result = labels.str.strip().str.replace(r"\s+", " ", regex=True).str.replace(r"\bPen\b", "Pencil", case=False, regex=True)
print(result)
assert result.iloc[0] == "Pencil blue"
assert result.iloc[1] == "Notebook A5"
assert pd.isna(result.iloc[2])

Expected output

0    Pencil blue
1    Notebook A5
2           <NA>
dtype: string

Practice in Google Colab 🔝

Open in Google Colab View on GitHub
Run the examples in order, change the replacement rules and compare raw and cleaned values. Save your own copy to keep edits. All sample tables are included in the notebook.

Continue with data cleaning, case conversion, string slicing and splitting text.

Reference: Pandas Series.str.replace documentation.




Subscribe to our YouTube Channel here



plus2net.com







Python Video Tutorials
Python SQLite Video Tutorials
Python MySQL Video Tutorials
Python Tkinter Video Tutorials
✖
We use cookies to improve your browsing experience. . Learn more
HTML MySQL PHP JavaScript ASP Photoshop Articles Contact us
© 2000-2026 plus2net.com All rights reserved worldwide Privacy Policy Disclaimer