Pandas str.zfill(): Pad and Validate Identifiers

Use str.zfill(width) to add leading zeros until text reaches a minimum width. Longer values remain unchanged, and letters can be padded too. Validate identifiers before formatting and preserve missing values.

Pad a numeric column after converting to text 🔝

zfill(width) sets a minimum string length. Convert the numeric column deliberately and keep a separate numeric copy if it is needed for calculations. See astype() and dtypes.

import pandas as pd
df = pd.DataFrame({"NAME": ["Ravi", "Raju", "Alex"], "ID": [1, 2, 3], "MATH": [30, 40, 50], "ENGLISH": [20, 30, 40]})
df["MATH"] = df["MATH"].astype("string").str.zfill(3)
print(df)
assert df["MATH"].tolist() == ["030", "040", "050"]

Expected output

   NAME  ID MATH  ENGLISH
0  Ravi   1  030       20
1  Raju   2  040       30
2  Alex   3  050       40

Width is a minimum, not a maximum 🔝

Strings already at or above the requested width are unchanged. Empty strings are padded too: they do not become missing automatically. Reject overlong identifiers separately instead of assuming zfill() truncates them.

values = pd.Series(["7", "58", "105", "6000", ""], dtype="string")
print(values.str.zfill(3))
assert values.str.zfill(3).tolist() == ["007", "058", "105", "6000", "000"]

Expected output

0     007
1     058
2     105
3    6000
4     000
dtype: string

Inspect signs and mixed values 🔝

Padding is string formatting, not numeric validation. This example includes signed values, letters and a value wider than three characters. The displayed results are generated by the current examples; exclude signs when your identifier schema permits digits only.

df = pd.DataFrame({"NAME": ["Ravi", "Raju", "Alex", "King", "Queen"], "ID": [1, 2, 3, 4, 5], "MATH": [-3, -40, 5, "abc", 6000], "ENGLISH": [20, 30, 40, 50, 10]})
df["MATH"] = df["MATH"].astype("string").str.zfill(3)
print(df)
assert df["MATH"].tolist() == ["-03", "-40", "005", "abc", "6000"]

Expected output

    NAME  ID  MATH  ENGLISH
0   Ravi   1   -03       20
1   Raju   2   -40       30
2   Alex   3   005       40
3   King   4   abc       50
4  Queen   5  6000       10

Format product codes to five characters 🔝

Use a known width from the receiving system or data specification. Padding is useful for identifiers, but is not a substitute for numeric display formats for amounts.

products = pd.DataFrame({"Product": ["A", "B", "C"], "Code": [7, 58, 105]})
products["Code"] = products["Code"].astype("string").str.zfill(5)
print(products)
assert products["Code"].tolist() == ["00007", "00058", "00105"]

Expected output

  Product   Code
0       A  00007
1       B  00058
2       C  00105

Letters are padded as well 🔝

zfill() does not distinguish numeric strings from alphabetic or alphanumeric strings. Short non-numeric values also receive zeros. Decide whether that transformation is allowed before applying it.

mixed = pd.DataFrame({"Category": ["alpha", "beta", "gamma"], "ID": [5, "x12", "abc"]})
mixed["ID"] = mixed["ID"].astype("string").str.zfill(4)
print(mixed)
codes = pd.Series(["A1", "B12", "C3"], dtype="string")
print(codes.str.zfill(4))
assert codes.str.zfill(4).tolist() == ["00A1", "0B12", "00C3"]

Expected output

  Category    ID
0    alpha  0005
1     beta  0x12
2    gamma  0abc
0    00A1
1    0B12
2    00C3
dtype: string

Preserve missing values with nullable string conversion 🔝

Use astype("string") when converting mixed input and preserving missing entries matters. Do not fill missing IDs with zero: zero can be a valid identifier. Floating-point input can stringify as 7.0 rather than 7 and needs validation before padding.

values = pd.Series([7, None, "12"], dtype="object")
text = values.astype("string")
print(text.str.zfill(4))
assert text.str.zfill(4).iloc[0] == "0007"
assert pd.isna(text.str.zfill(4).iloc[1])

Expected output

0    0007
1    <NA>
2    0012
dtype: string

Trim and validate before padding 🔝

Whitespace counts toward string length and is retained unless removed. Convert blank text to missing or reject it according to your rules. Use str.len() to inspect lengths.

values = pd.Series([" 7 ", "12", "  ", None], dtype="string")
trimmed = values.str.strip()
trimmed = trimmed.mask(trimmed.eq(""), pd.NA)
print(trimmed.str.zfill(4))
assert trimmed.str.zfill(4).iloc[0] == "0007"
assert trimmed.isna().sum() == 2

Expected output

0    0007
1    0012
2    <NA>
3    <NA>
dtype: string

Practical workflow: standardise a four-digit identifier 🔝

This synthetic schema allows one to four ASCII digits and pads them to four characters. Reject missing, blank, signed, decimal, alphabetic and overlong values before padding. Keep original values and rejected rows for review.

raw = pd.DataFrame({"id": ["7", "0058", " 105 ", "", None, "12345", "-3", "7.0", "A12"]})
text = raw["id"].astype("string").str.strip()
valid = text.str.fullmatch(r"[0-9]{1,4}", na=False)
accepted = raw.loc[valid].copy()
accepted["normalised_id"] = text.loc[valid].str.zfill(4)
review = raw.loc[~valid].copy()
print(accepted)
print("Review rows:")
print(review)
assert accepted["normalised_id"].tolist() == ["0007", "0058", "0105"]

Expected output

      id normalised_id
0      7          0007
1   0058          0058
2   105           0105
Review rows:
      id
3       
4    NaN
5  12345
6     -3
7    7.0
8    A12

Check whether normalisation creates duplicate keys 🔝

Different raw spellings can become the same padded identifier. That may be intentional equivalence or a data conflict. Review collisions before joining tables or removing records; see duplicate handling.

source = pd.DataFrame({"id": pd.Series(["7", "007", "8"], dtype="string")})
source["key"] = source["id"].str.zfill(4)
duplicates = source.duplicated("key", keep=False)
print(source.loc[duplicates])
assert duplicates.tolist() == [True, True, False]

Expected output

    id   key
0    7  0007
1  007  0007

Read identifiers as text and preserve them on export 🔝

Specify the identifier dtype when reading CSV data. A CSV does not store a column dtype, so spreadsheet software may reinterpret padded codes as numbers. Check the receiving application and keep identifiers as text during import.

from io import StringIO
csv_source = "id,product\n0007,Pen\n0058,Notebook\n"
imported = pd.read_csv(StringIO(csv_source), dtype={"id": "string"})
print(imported)
assert imported["id"].tolist() == ["0007", "0058"]
print(accepted.to_csv(index=False).rstrip())
# Optional Colab export:
# accepted.to_csv("padded_identifiers.csv", index=False)

Expected output

     id   product
0  0007       Pen
1  0058  Notebook
id,normalised_id
7,0007
0058,0058
 105 ,0105

Compare padding methods 🔝

rjust() and pad() also support left padding; ljust() pads the right. The comparison uses unsigned text because sign handling deserves separate testing. If a code has a letter prefix plus a numeric suffix, validate and separate those fields before padding just the suffix; see slicing, splitting and recombining fields.

s = pd.Series(["7", "58"], dtype="string")
print(pd.DataFrame({"zfill": s.str.zfill(4), "rjust": s.str.rjust(4, fillchar="0"), "right_padding": s.str.ljust(4, fillchar="0")}))
assert s.str.zfill(4).equals(s.str.rjust(4, fillchar="0"))

Expected output

  zfill rjust right_padding
0  0007  0007          7000
1  0058  0058          5800

Common questions 🔝

Why did letters receive zeros? zfill formats strings without checking whether they are numbers. Why is the result wider than requested? Width is a minimum. Can padding restore a lost original ID? Only when a known schema defines its width; it cannot infer an unknown identifier. Why is a blank now all zeros? Empty text is a string, not missing. Does padding change the source? Assign the result explicitly to store it.

Exercise: validate a five-digit code 🔝

Create values 9, 012, 123456, a blank and a missing value as text. Accept only one to five ASCII digits, pad accepted values to width five and retain rejected values. Expect two accepted codes.

Exercise solution 🔝

Validate length and allowed characters before applying padding.

practice = pd.Series(["9", "012", "123456", "", None], dtype="string")
ok = practice.str.fullmatch(r"[0-9]{1,5}", na=False)
result = practice.loc[ok].str.zfill(5)
print(result)
print("Review values:")
print(practice.loc[~ok])
assert result.tolist() == ["00009", "00012"]

Expected output

0    00009
1    00012
dtype: string
Review values:
2    123456
3          
4      <NA>
dtype: string

Practice in Google Colab 🔝

Open in Google Colab View on GitHub
Run the examples in order, change identifier widths and invalid inputs, then compare accepted values and review rows. Save your own copy to keep edits. All sample tables are included in the notebook.

Continue with the Data Cleaning and String Operations hub, string lengths, CSV input and Excel export.

Reference: Pandas Series.str.zfill documentation.




Subscribe to our YouTube Channel here



plus2net.com







Python Video Tutorials
Python SQLite Video Tutorials
Python MySQL Video Tutorials
Python Tkinter Video Tutorials
✖
We use cookies to improve your browsing experience. . Learn more
HTML MySQL PHP JavaScript ASP Photoshop Articles Contact us
© 2000-2026 plus2net.com All rights reserved worldwide Privacy Policy Disclaimer