describe() returns summary statistics for a Series or DataFrame. Numeric data produces counts, mean, standard deviation, minimum, quartiles and maximum. Text and categories produce counts, distinct values and frequency summaries.
Use the report to understand your data before cleaning it, and compare it afterwards. Statistics alone do not explain why a value is missing or whether a replacement is appropriate.
DataFrame.describe(percentiles=None, include=None, exclude=None)
Percentiles are fractions between 0 and 1; include and exclude select dtypes. Outputs below were checked locally with Pandas 3.0.1. Formatting and tied top values may differ in your environment.
Show Table of Contentsdescribe() summarises non-missing values. Run this setup example first; later student examples call student_data() to start from fresh data.
import pandas as pd
def student_data():
df = pd.DataFrame({
"NAME": ["Ravi", "Raju", "Alex", "Ron", "King", "Jack"],
"ID": [1, 2, 3, 4, 5, 6],
"MATH": [80, 40, 70, 70, 60, 30],
"ENGLISH": [80, 70, 40, 50, 60, 30],
})
# Explicit dtype makes text selection consistent across Pandas versions.
df["NAME"] = df["NAME"].astype("string")
return df
students = student_data()
print(students["MATH"].describe().to_string())
Expected output
count 6.000000
mean 58.333333
std 19.407902
min 30.000000
25% 45.000000
50% 65.000000
75% 70.000000
max 80.000000
The mean mark is about 58.33, while the median is 65. These are different summaries of the same six marks.
For mixed numeric and text data, the default describes numeric columns. A numeric identifier is included even though its average usually has no useful meaning.
students = student_data()
print(students.describe().to_string())
Expected output
ID MATH ENGLISH
count 6.000000 6.000000 6.000000
mean 3.500000 58.333333 55.000000
std 1.870829 19.407902 18.708287
min 1.000000 30.000000 30.000000
25% 2.250000 45.000000 42.500000
50% 3.500000 65.000000 55.000000
75% 4.750000 70.000000 67.500000
max 6.000000 80.000000 80.000000
For a meaningful marks report, select students[["MATH", "ENGLISH"]] before describing the data. Missing values are excluded independently in each column.
| Statistic | Interpretation |
|---|---|
| count | Number of non-missing observations in that column, not necessarily the number of rows. |
| mean | Arithmetic average. Large or small unusual values can affect it strongly. |
| std | Sample standard deviation, using a denominator of count minus one. With only one observation it is NaN. |
| min / max | Smallest and largest observed values. |
| 25% / 50% / 75% | Quartiles; 50% is the median. Interpolation can produce a value that does not appear in the input. |
| unique | Number of distinct non-missing text or category values. |
| top / freq | A most frequent value and its frequency. In a tie, do not rely on which value is selected as top. |
Percentiles use fractions from 0 to 1. The 50% median is included even when it is not in the supplied list.
students = student_data()
print(students["MATH"].describe(percentiles=[0.45, 0.68, 0.89]).to_string())
Expected output
count 6.000000
mean 58.333333
std 19.407902
min 30.000000
45% 62.500000
68% 70.000000
89% 74.500000
max 80.000000
include="all" combines statistics appropriate for each dtype. NaN in the combined report can mean a statistic does not apply, rather than a missing source record.
students = student_data()
print(students.describe(include="all").to_string())
Expected output
NAME ID MATH ENGLISH
count 6 6.000000 6.000000 6.000000
unique 6 NaN NaN NaN
top Ravi NaN NaN NaN
freq 1 NaN NaN NaN
mean NaN 3.500000 58.333333 55.000000
std NaN 1.870829 19.407902 18.708287
min NaN 1.000000 30.000000 30.000000
25% NaN 2.250000 45.000000 42.500000
50% NaN 3.500000 65.000000 55.000000
75% NaN 4.750000 70.000000 67.500000
max NaN 6.000000 80.000000 80.000000
All student names occur once. Any of them may appear as top; the important result is freq = 1.
Use an explicit string dtype and include=["string"]. For columns genuinely stored as object dtype, use include=["object"]. Avoid the removed NumPy alias np.object.
students = student_data()
print(students.describe(include=["string"]).to_string())
Expected output
NAME
count 6
unique 6
top Ravi
freq 1
An object column can contain text, but it is different from an explicit string column. This example deliberately converts the name column to object dtype.
students = student_data()
students["NAME"] = students["NAME"].astype("object")
print(students.describe(include=["object"]).to_string())
Expected output
NAME
count 6
unique 6
top Ravi
freq 1
include=["number"] selects numeric dtypes without requiring a NumPy import. It serves the same selection purpose as include=[np.number].
students = student_data()
print(students.describe(include=["number"]).to_string())
Expected output
ID MATH ENGLISH
count 6.000000 6.000000 6.000000
mean 3.500000 58.333333 55.000000
std 1.870829 19.407902 18.708287
min 1.000000 30.000000 30.000000
25% 2.250000 45.000000 42.500000
50% 3.500000 65.000000 55.000000
75% 4.750000 70.000000 67.500000
max 6.000000 80.000000 80.000000
Excluding categories retains the remaining eligible column types. Here there is no category column, so none is removed.
students = student_data()
print(students.describe(exclude=["category"]).to_string())
Expected output
NAME ID MATH ENGLISH
count 6 6.000000 6.000000 6.000000
unique 6 NaN NaN NaN
top Ravi NaN NaN NaN
freq 1 NaN NaN NaN
mean NaN 3.500000 58.333333 55.000000
std NaN 1.870829 19.407902 18.708287
min NaN 1.000000 30.000000 30.000000
25% NaN 2.250000 45.000000 42.500000
50% NaN 3.500000 65.000000 55.000000
75% NaN 4.750000 70.000000 67.500000
max NaN 6.000000 80.000000 80.000000
Exclude numeric columns to focus on the names.
students = student_data()
print(students.describe(exclude=["number"]).to_string())
Expected output
NAME
count 6
unique 6
top Ravi
freq 1
Exclude the explicit string dtype to leave numeric statistics. For object-dtype text, specify exclude=["object"] instead.
students = student_data()
print(students.describe(exclude=["string"]).to_string())
Expected output
ID MATH ENGLISH
count 6.000000 6.000000 6.000000
mean 3.500000 58.333333 55.000000
std 1.870829 19.407902 18.708287
min 1.000000 30.000000 30.000000
25% 2.250000 45.000000 42.500000
50% 3.500000 65.000000 55.000000
75% 4.750000 70.000000 67.500000
max 6.000000 80.000000 80.000000
A category column is useful for repeated labels such as grades. This example retains the student-grade lesson and creates a fresh DataFrame.
students = student_data()
students["grade"] = pd.Series(["a", "c", "b", "b", "b", "c"], dtype="category")
print(students.describe(include="all").to_string())
Expected output
NAME ID MATH ENGLISH grade
count 6 6.000000 6.000000 6.000000 6
unique 6 NaN NaN NaN 3
top Ravi NaN NaN NaN b
freq 1 NaN NaN NaN 3
mean NaN 3.500000 58.333333 55.000000 NaN
std NaN 1.870829 19.407902 18.708287 NaN
min NaN 1.000000 30.000000 30.000000 NaN
25% NaN 2.250000 45.000000 42.500000 NaN
50% NaN 3.500000 65.000000 55.000000 NaN
75% NaN 4.750000 70.000000 67.500000 NaN
max NaN 6.000000 80.000000 80.000000 NaN
There are three distinct observed grades. Grade b appears three times.
students = student_data()
students["grade"] = pd.Series(["a", "c", "b", "b", "b", "c"], dtype="category")
print(students.describe(include=["category"]).to_string())
Expected output
grade
count 6
unique 3
top b
freq 3
The report below retains names and numeric values while excluding grades.
students = student_data()
students["grade"] = pd.Series(["a", "c", "b", "b", "b", "c"], dtype="category")
print(students.describe(exclude=["category"]).to_string())
Expected output
NAME ID MATH ENGLISH
count 6 6.000000 6.000000 6.000000
unique 6 NaN NaN NaN
top Ravi NaN NaN NaN
freq 1 NaN NaN NaN
mean NaN 3.500000 58.333333 55.000000
std NaN 1.870829 19.407902 18.708287
min NaN 1.000000 30.000000 30.000000
25% NaN 2.250000 45.000000 42.500000
50% NaN 3.500000 65.000000 55.000000
75% NaN 4.750000 70.000000 67.500000
max NaN 6.000000 80.000000 80.000000
Count shows only available observations. An all-missing numeric column has a count of zero and undefined summary statistics.
values = pd.DataFrame({"observed": [10.0, None, 30.0], "empty": [float("nan")] * 3})
print(values.describe().to_string())
print("Missing values:")
print(values.isna().sum().to_string())
Expected output
observed empty
count 2.000000 0.0
mean 20.000000 NaN
std 14.142136 NaN
min 10.000000 NaN
25% 15.000000 NaN
50% 20.000000 NaN
75% 25.000000 NaN
max 30.000000 NaN
Missing values:
observed 1
empty 3
Use isnull() or isna() for missingness. Describe does not tell you which rows are missing.
Zero filling changes the average and count. Median filling estimates a value and changes count too, even when the average stays the same.
scores = pd.Series([10.0, 20.0, None, 30.0])
report = pd.DataFrame({"observed": scores.describe(), "zero_filled": scores.fillna(0).describe(), "median_filled": scores.fillna(scores.median()).describe()})
print(report.loc[["count", "mean", "50%", "std"]].to_string())
Expected output
observed zero_filled median_filled
count 3.0 4.000000 4.000000
mean 20.0 15.000000 20.000000
50% 20.0 15.000000 20.000000
std 10.0 12.909944 8.164966
A higher count after filling is not evidence that you collected more observations. Read fillna() and practical missing-value handling before choosing a replacement.
Use the same synthetic five-order CSV as the fillna lesson. The code loads it from Plus2Net; an internet connection is required. Exclude order IDs from the numeric report.
import pandas as pd
DATA_URL = "https://www.plus2net.com/python/download/pandas_fillna_sample_sales.csv"
sales = pd.read_csv(DATA_URL, dtype={"category": "string"})
print(sales[["quantity", "unit_price", "discount"]].describe().to_string())
print("\nMissing values:")
print(sales[["quantity", "unit_price", "discount"]].isna().sum().to_string())
print("\nRows:", len(sales))
Expected output
quantity unit_price discount
count 4.000000 4.000000 3.000000
mean 2.000000 18.750000 1.666667
std 0.816497 8.539126 2.886751
min 1.000000 10.000000 0.000000
25% 1.750000 13.750000 0.000000
50% 2.000000 17.500000 0.000000
75% 2.250000 22.500000 2.500000
max 3.000000 30.000000 5.000000
Missing values:
quantity 1
unit_price 1
discount 2
Rows: 5
There are five rows but only four known quantities and prices, and three recorded discounts. That is a useful cleaning signal. The price median is not permission to invent the price of an incomplete order.
Continue from the preceding sales example. A blank discount means zero under this exercise rule; quantities and prices remain unknown. Compare the reports.
cleaned = sales.fillna({"category": "Unknown", "discount": 0})
print(cleaned[["quantity", "unit_price", "discount"]].describe().to_string())
assert len(cleaned) == len(sales)
assert cleaned["quantity"].isna().sum() == sales["quantity"].isna().sum()
assert cleaned["unit_price"].isna().sum() == sales["unit_price"].isna().sum()
print("Row count and unknown quantities/prices preserved.")
Expected output
quantity unit_price discount
count 4.000000 4.000000 5.000000
mean 2.000000 18.750000 1.000000
std 0.816497 8.539126 2.236068
min 1.000000 10.000000 0.000000
25% 1.750000 13.750000 0.000000
50% 2.000000 17.500000 0.000000
75% 2.250000 22.500000 0.000000
max 3.000000 30.000000 5.000000
Row count and unknown quantities/prices preserved.
The discount count becomes five and its mean changes. Quantity and price counts remain four. Always explain the rule behind a change in statistics.
Run the examples and exercises in the accompanying notebook. Sign in, choose File → Save a copy in Drive, then run the cells in order. The sales CSV loads automatically from Plus2Net.
Open in Google Colab View on GitHubUse describe to understand observed values, then investigate missingness and choose justified cleaning rules. Select meaningful measurements, interpret statistics according to dtype, and verify changes after cleaning.
Return to Data Exploration and Analysis, or continue with fillna, DataFrame dtypes, or the Pandas reporting exercises. The exercises connect aggregation, groupby and reporting; describe itself does not merge tables.
Reference: Official Pandas describe documentation.
Pandas Tutorials DataFrameAuthor & Instructor at plus2net
I write and maintain practical tutorials on Python, PHP, SQL, JavaScript, HTML, jQuery, and web development at plus2net. The tutorials focus on clear explanations, working examples, and code that readers can test and adapt while learning.