Use str.zfill(width) to add leading zeros until text reaches a minimum width. Longer values remain unchanged, and letters can be padded too. Validate identifiers before formatting and preserve missing values.
zfill(width) sets a minimum string length. Convert the numeric column deliberately and keep a separate numeric copy if it is needed for calculations. See astype() and dtypes.
import pandas as pd
df = pd.DataFrame({"NAME": ["Ravi", "Raju", "Alex"], "ID": [1, 2, 3], "MATH": [30, 40, 50], "ENGLISH": [20, 30, 40]})
df["MATH"] = df["MATH"].astype("string").str.zfill(3)
print(df)
assert df["MATH"].tolist() == ["030", "040", "050"]Expected output
NAME ID MATH ENGLISH
0 Ravi 1 030 20
1 Raju 2 040 30
2 Alex 3 050 40Strings already at or above the requested width are unchanged. Empty strings are padded too: they do not become missing automatically. Reject overlong identifiers separately instead of assuming zfill() truncates them.
values = pd.Series(["7", "58", "105", "6000", ""], dtype="string")
print(values.str.zfill(3))
assert values.str.zfill(3).tolist() == ["007", "058", "105", "6000", "000"]Expected output
0 007
1 058
2 105
3 6000
4 000
dtype: stringPadding is string formatting, not numeric validation. This example includes signed values, letters and a value wider than three characters. The displayed results are generated by the current examples; exclude signs when your identifier schema permits digits only.
df = pd.DataFrame({"NAME": ["Ravi", "Raju", "Alex", "King", "Queen"], "ID": [1, 2, 3, 4, 5], "MATH": [-3, -40, 5, "abc", 6000], "ENGLISH": [20, 30, 40, 50, 10]})
df["MATH"] = df["MATH"].astype("string").str.zfill(3)
print(df)
assert df["MATH"].tolist() == ["-03", "-40", "005", "abc", "6000"]Expected output
NAME ID MATH ENGLISH
0 Ravi 1 -03 20
1 Raju 2 -40 30
2 Alex 3 005 40
3 King 4 abc 50
4 Queen 5 6000 10Use a known width from the receiving system or data specification. Padding is useful for identifiers, but is not a substitute for numeric display formats for amounts.
products = pd.DataFrame({"Product": ["A", "B", "C"], "Code": [7, 58, 105]})
products["Code"] = products["Code"].astype("string").str.zfill(5)
print(products)
assert products["Code"].tolist() == ["00007", "00058", "00105"]Expected output
Product Code
0 A 00007
1 B 00058
2 C 00105zfill() does not distinguish numeric strings from alphabetic or alphanumeric strings. Short non-numeric values also receive zeros. Decide whether that transformation is allowed before applying it.
mixed = pd.DataFrame({"Category": ["alpha", "beta", "gamma"], "ID": [5, "x12", "abc"]})
mixed["ID"] = mixed["ID"].astype("string").str.zfill(4)
print(mixed)
codes = pd.Series(["A1", "B12", "C3"], dtype="string")
print(codes.str.zfill(4))
assert codes.str.zfill(4).tolist() == ["00A1", "0B12", "00C3"]Expected output
Category ID
0 alpha 0005
1 beta 0x12
2 gamma 0abc
0 00A1
1 0B12
2 00C3
dtype: stringUse astype("string") when converting mixed input and preserving missing entries matters. Do not fill missing IDs with zero: zero can be a valid identifier. Floating-point input can stringify as 7.0 rather than 7 and needs validation before padding.
values = pd.Series([7, None, "12"], dtype="object")
text = values.astype("string")
print(text.str.zfill(4))
assert text.str.zfill(4).iloc[0] == "0007"
assert pd.isna(text.str.zfill(4).iloc[1])Expected output
0 0007
1 <NA>
2 0012
dtype: stringWhitespace counts toward string length and is retained unless removed. Convert blank text to missing or reject it according to your rules. Use str.len() to inspect lengths.
values = pd.Series([" 7 ", "12", " ", None], dtype="string")
trimmed = values.str.strip()
trimmed = trimmed.mask(trimmed.eq(""), pd.NA)
print(trimmed.str.zfill(4))
assert trimmed.str.zfill(4).iloc[0] == "0007"
assert trimmed.isna().sum() == 2Expected output
0 0007
1 0012
2 <NA>
3 <NA>
dtype: stringThis synthetic schema allows one to four ASCII digits and pads them to four characters. Reject missing, blank, signed, decimal, alphabetic and overlong values before padding. Keep original values and rejected rows for review.
raw = pd.DataFrame({"id": ["7", "0058", " 105 ", "", None, "12345", "-3", "7.0", "A12"]})
text = raw["id"].astype("string").str.strip()
valid = text.str.fullmatch(r"[0-9]{1,4}", na=False)
accepted = raw.loc[valid].copy()
accepted["normalised_id"] = text.loc[valid].str.zfill(4)
review = raw.loc[~valid].copy()
print(accepted)
print("Review rows:")
print(review)
assert accepted["normalised_id"].tolist() == ["0007", "0058", "0105"]Expected output
id normalised_id
0 7 0007
1 0058 0058
2 105 0105
Review rows:
id
3
4 NaN
5 12345
6 -3
7 7.0
8 A12Different raw spellings can become the same padded identifier. That may be intentional equivalence or a data conflict. Review collisions before joining tables or removing records; see duplicate handling.
source = pd.DataFrame({"id": pd.Series(["7", "007", "8"], dtype="string")})
source["key"] = source["id"].str.zfill(4)
duplicates = source.duplicated("key", keep=False)
print(source.loc[duplicates])
assert duplicates.tolist() == [True, True, False]Expected output
id key
0 7 0007
1 007 0007Specify the identifier dtype when reading CSV data. A CSV does not store a column dtype, so spreadsheet software may reinterpret padded codes as numbers. Check the receiving application and keep identifiers as text during import.
from io import StringIO
csv_source = "id,product\n0007,Pen\n0058,Notebook\n"
imported = pd.read_csv(StringIO(csv_source), dtype={"id": "string"})
print(imported)
assert imported["id"].tolist() == ["0007", "0058"]
print(accepted.to_csv(index=False).rstrip())
# Optional Colab export:
# accepted.to_csv("padded_identifiers.csv", index=False)Expected output
id product
0 0007 Pen
1 0058 Notebook
id,normalised_id
7,0007
0058,0058
105 ,0105rjust() and pad() also support left padding; ljust() pads the right. The comparison uses unsigned text because sign handling deserves separate testing. If a code has a letter prefix plus a numeric suffix, validate and separate those fields before padding just the suffix; see slicing, splitting and recombining fields.
s = pd.Series(["7", "58"], dtype="string")
print(pd.DataFrame({"zfill": s.str.zfill(4), "rjust": s.str.rjust(4, fillchar="0"), "right_padding": s.str.ljust(4, fillchar="0")}))
assert s.str.zfill(4).equals(s.str.rjust(4, fillchar="0"))Expected output
zfill rjust right_padding
0 0007 0007 7000
1 0058 0058 5800Why did letters receive zeros? zfill formats strings without checking whether they are numbers. Why is the result wider than requested? Width is a minimum. Can padding restore a lost original ID? Only when a known schema defines its width; it cannot infer an unknown identifier. Why is a blank now all zeros? Empty text is a string, not missing. Does padding change the source? Assign the result explicitly to store it.
Create values 9, 012, 123456, a blank and a missing value as text. Accept only one to five ASCII digits, pad accepted values to width five and retain rejected values. Expect two accepted codes.
Validate length and allowed characters before applying padding.
practice = pd.Series(["9", "012", "123456", "", None], dtype="string")
ok = practice.str.fullmatch(r"[0-9]{1,5}", na=False)
result = practice.loc[ok].str.zfill(5)
print(result)
print("Review values:")
print(practice.loc[~ok])
assert result.tolist() == ["00009", "00012"]Expected output
0 00009
1 00012
dtype: string
Review values:
2 123456
3
4 <NA>
dtype: stringOpen in Google Colab View on GitHub
Run the examples in order, change identifier widths and invalid inputs, then compare accepted values and review rows. Save your own copy to keep edits. All sample tables are included in the notebook.
Continue with the Data Cleaning and String Operations hub, string lengths, CSV input and Excel export.
Reference: Pandas Series.str.zfill documentation.
Author & Instructor at plus2net
I write and maintain practical tutorials on Python, PHP, SQL, JavaScript, HTML, jQuery, and web development at plus2net. The tutorials focus on clear explanations, working examples, and code that readers can test and adapt while learning.