Pandas describe(): Understand Summary Statistics

describe() returns summary statistics for a Series or DataFrame. Numeric data produces counts, mean, standard deviation, minimum, quartiles and maximum. Text and categories produce counts, distinct values and frequency summaries.

Use the report to understand your data before cleaning it, and compare it afterwards. Statistics alone do not explain why a value is missing or whether a replacement is appropriate.

DataFrame.describe(percentiles=None, include=None, exclude=None)

Percentiles are fractions between 0 and 1; include and exclude select dtypes. Outputs below were checked locally with Pandas 3.0.1. Formatting and tied top values may differ in your environment.


Describe One Numeric Column

describe() summarises non-missing values. Run this setup example first; later student examples call student_data() to start from fresh data.

import pandas as pd

def student_data():
    df = pd.DataFrame({
        "NAME": ["Ravi", "Raju", "Alex", "Ron", "King", "Jack"],
        "ID": [1, 2, 3, 4, 5, 6],
        "MATH": [80, 40, 70, 70, 60, 30],
        "ENGLISH": [80, 70, 40, 50, 60, 30],
    })
    # Explicit dtype makes text selection consistent across Pandas versions.
    df["NAME"] = df["NAME"].astype("string")
    return df

students = student_data()
print(students["MATH"].describe().to_string())

Expected output

count     6.000000
mean     58.333333
std      19.407902
min      30.000000
25%      45.000000
50%      65.000000
75%      70.000000
max      80.000000

The mean mark is about 58.33, while the median is 65. These are different summaries of the same six marks.

Describe a DataFrame

For mixed numeric and text data, the default describes numeric columns. A numeric identifier is included even though its average usually has no useful meaning.

students = student_data()
print(students.describe().to_string())

Expected output

             ID       MATH    ENGLISH
count  6.000000   6.000000   6.000000
mean   3.500000  58.333333  55.000000
std    1.870829  19.407902  18.708287
min    1.000000  30.000000  30.000000
25%    2.250000  45.000000  42.500000
50%    3.500000  65.000000  55.000000
75%    4.750000  70.000000  67.500000
max    6.000000  80.000000  80.000000

For a meaningful marks report, select students[["MATH", "ENGLISH"]] before describing the data. Missing values are excluded independently in each column.

Understand the Statistics

StatisticInterpretation
countNumber of non-missing observations in that column, not necessarily the number of rows.
meanArithmetic average. Large or small unusual values can affect it strongly.
stdSample standard deviation, using a denominator of count minus one. With only one observation it is NaN.
min / maxSmallest and largest observed values.
25% / 50% / 75%Quartiles; 50% is the median. Interpolation can produce a value that does not appear in the input.
uniqueNumber of distinct non-missing text or category values.
top / freqA most frequent value and its frequency. In a tie, do not rely on which value is selected as top.

Choose Custom Percentiles

Percentiles use fractions from 0 to 1. The 50% median is included even when it is not in the supplied list.

students = student_data()
print(students["MATH"].describe(percentiles=[0.45, 0.68, 0.89]).to_string())

Expected output

count     6.000000
mean     58.333333
std      19.407902
min      30.000000
45%      62.500000
68%      70.000000
89%      74.500000
max      80.000000

Include All Column Types

include="all" combines statistics appropriate for each dtype. NaN in the combined report can mean a statistic does not apply, rather than a missing source record.

students = student_data()
print(students.describe(include="all").to_string())

Expected output

        NAME        ID       MATH    ENGLISH
count      6  6.000000   6.000000   6.000000
unique     6       NaN        NaN        NaN
top     Ravi       NaN        NaN        NaN
freq       1       NaN        NaN        NaN
mean     NaN  3.500000  58.333333  55.000000
std      NaN  1.870829  19.407902  18.708287
min      NaN  1.000000  30.000000  30.000000
25%      NaN  2.250000  45.000000  42.500000
50%      NaN  3.500000  65.000000  55.000000
75%      NaN  4.750000  70.000000  67.500000
max      NaN  6.000000  80.000000  80.000000

All student names occur once. Any of them may appear as top; the important result is freq = 1.

Select Text Columns

Use an explicit string dtype and include=["string"]. For columns genuinely stored as object dtype, use include=["object"]. Avoid the removed NumPy alias np.object.

students = student_data()
print(students.describe(include=["string"]).to_string())

Expected output

        NAME
count      6
unique     6
top     Ravi
freq       1

Describe Object Dtype Columns

An object column can contain text, but it is different from an explicit string column. This example deliberately converts the name column to object dtype.

students = student_data()
students["NAME"] = students["NAME"].astype("object")
print(students.describe(include=["object"]).to_string())

Expected output

        NAME
count      6
unique     6
top     Ravi
freq       1

Select Numeric Columns

include=["number"] selects numeric dtypes without requiring a NumPy import. It serves the same selection purpose as include=[np.number].

students = student_data()
print(students.describe(include=["number"]).to_string())

Expected output

             ID       MATH    ENGLISH
count  6.000000   6.000000   6.000000
mean   3.500000  58.333333  55.000000
std    1.870829  19.407902  18.708287
min    1.000000  30.000000  30.000000
25%    2.250000  45.000000  42.500000
50%    3.500000  65.000000  55.000000
75%    4.750000  70.000000  67.500000
max    6.000000  80.000000  80.000000

Exclude a Dtype

Excluding categories retains the remaining eligible column types. Here there is no category column, so none is removed.

students = student_data()
print(students.describe(exclude=["category"]).to_string())

Expected output

        NAME        ID       MATH    ENGLISH
count      6  6.000000   6.000000   6.000000
unique     6       NaN        NaN        NaN
top     Ravi       NaN        NaN        NaN
freq       1       NaN        NaN        NaN
mean     NaN  3.500000  58.333333  55.000000
std      NaN  1.870829  19.407902  18.708287
min      NaN  1.000000  30.000000  30.000000
25%      NaN  2.250000  45.000000  42.500000
50%      NaN  3.500000  65.000000  55.000000
75%      NaN  4.750000  70.000000  67.500000
max      NaN  6.000000  80.000000  80.000000

Exclude Numeric Columns

Exclude numeric columns to focus on the names.

students = student_data()
print(students.describe(exclude=["number"]).to_string())

Expected output

        NAME
count      6
unique     6
top     Ravi
freq       1

Exclude Text Columns

Exclude the explicit string dtype to leave numeric statistics. For object-dtype text, specify exclude=["object"] instead.

students = student_data()
print(students.describe(exclude=["string"]).to_string())

Expected output

             ID       MATH    ENGLISH
count  6.000000   6.000000   6.000000
mean   3.500000  58.333333  55.000000
std    1.870829  19.407902  18.708287
min    1.000000  30.000000  30.000000
25%    2.250000  45.000000  42.500000
50%    3.500000  65.000000  55.000000
75%    4.750000  70.000000  67.500000
max    6.000000  80.000000  80.000000

Create and Describe Category Data

A category column is useful for repeated labels such as grades. This example retains the student-grade lesson and creates a fresh DataFrame.

students = student_data()
students["grade"] = pd.Series(["a", "c", "b", "b", "b", "c"], dtype="category")
print(students.describe(include="all").to_string())

Expected output

        NAME        ID       MATH    ENGLISH grade
count      6  6.000000   6.000000   6.000000     6
unique     6       NaN        NaN        NaN     3
top     Ravi       NaN        NaN        NaN     b
freq       1       NaN        NaN        NaN     3
mean     NaN  3.500000  58.333333  55.000000   NaN
std      NaN  1.870829  19.407902  18.708287   NaN
min      NaN  1.000000  30.000000  30.000000   NaN
25%      NaN  2.250000  45.000000  42.500000   NaN
50%      NaN  3.500000  65.000000  55.000000   NaN
75%      NaN  4.750000  70.000000  67.500000   NaN
max      NaN  6.000000  80.000000  80.000000   NaN

Describe Only Categories

There are three distinct observed grades. Grade b appears three times.

students = student_data()
students["grade"] = pd.Series(["a", "c", "b", "b", "b", "c"], dtype="category")
print(students.describe(include=["category"]).to_string())

Expected output

       grade
count      6
unique     3
top        b
freq       3

Exclude the Category Column

The report below retains names and numeric values while excluding grades.

students = student_data()
students["grade"] = pd.Series(["a", "c", "b", "b", "b", "c"], dtype="category")
print(students.describe(exclude=["category"]).to_string())

Expected output

        NAME        ID       MATH    ENGLISH
count      6  6.000000   6.000000   6.000000
unique     6       NaN        NaN        NaN
top     Ravi       NaN        NaN        NaN
freq       1       NaN        NaN        NaN
mean     NaN  3.500000  58.333333  55.000000
std      NaN  1.870829  19.407902  18.708287
min      NaN  1.000000  30.000000  30.000000
25%      NaN  2.250000  45.000000  42.500000
50%      NaN  3.500000  65.000000  55.000000
75%      NaN  4.750000  70.000000  67.500000
max      NaN  6.000000  80.000000  80.000000

Missing Values and Small Samples

Count shows only available observations. An all-missing numeric column has a count of zero and undefined summary statistics.

values = pd.DataFrame({"observed": [10.0, None, 30.0], "empty": [float("nan")] * 3})
print(values.describe().to_string())
print("Missing values:")
print(values.isna().sum().to_string())

Expected output

        observed  empty
count   2.000000    0.0
mean   20.000000    NaN
std    14.142136    NaN
min    10.000000    NaN
25%    15.000000    NaN
50%    20.000000    NaN
75%    25.000000    NaN
max    30.000000    NaN
Missing values:
observed    1
empty       3

Use isnull() or isna() for missingness. Describe does not tell you which rows are missing.

Compare Statistics Before and After Filling

Zero filling changes the average and count. Median filling estimates a value and changes count too, even when the average stays the same.

scores = pd.Series([10.0, 20.0, None, 30.0])
report = pd.DataFrame({"observed": scores.describe(), "zero_filled": scores.fillna(0).describe(), "median_filled": scores.fillna(scores.median()).describe()})
print(report.loc[["count", "mean", "50%", "std"]].to_string())

Expected output

       observed  zero_filled  median_filled
count       3.0     4.000000       4.000000
mean       20.0    15.000000      20.000000
50%        20.0    15.000000      20.000000
std        10.0    12.909944       8.164966

A higher count after filling is not evidence that you collected more observations. Read fillna() and practical missing-value handling before choosing a replacement.

Practical Example: Inspect a Sales CSV

Use the same synthetic five-order CSV as the fillna lesson. The code loads it from Plus2Net; an internet connection is required. Exclude order IDs from the numeric report.

import pandas as pd

DATA_URL = "https://www.plus2net.com/python/download/pandas_fillna_sample_sales.csv"
sales = pd.read_csv(DATA_URL, dtype={"category": "string"})
print(sales[["quantity", "unit_price", "discount"]].describe().to_string())
print("\nMissing values:")
print(sales[["quantity", "unit_price", "discount"]].isna().sum().to_string())
print("\nRows:", len(sales))

Expected output

       quantity  unit_price  discount
count  4.000000    4.000000  3.000000
mean   2.000000   18.750000  1.666667
std    0.816497    8.539126  2.886751
min    1.000000   10.000000  0.000000
25%    1.750000   13.750000  0.000000
50%    2.000000   17.500000  0.000000
75%    2.250000   22.500000  2.500000
max    3.000000   30.000000  5.000000

Missing values:
quantity      1
unit_price    1
discount      2

Rows: 5

There are five rows but only four known quantities and prices, and three recorded discounts. That is a useful cleaning signal. The price median is not permission to invent the price of an incomplete order.

Verify the Agreed Cleaning Rules

Continue from the preceding sales example. A blank discount means zero under this exercise rule; quantities and prices remain unknown. Compare the reports.

cleaned = sales.fillna({"category": "Unknown", "discount": 0})
print(cleaned[["quantity", "unit_price", "discount"]].describe().to_string())
assert len(cleaned) == len(sales)
assert cleaned["quantity"].isna().sum() == sales["quantity"].isna().sum()
assert cleaned["unit_price"].isna().sum() == sales["unit_price"].isna().sum()
print("Row count and unknown quantities/prices preserved.")

Expected output

       quantity  unit_price  discount
count  4.000000    4.000000  5.000000
mean   2.000000   18.750000  1.000000
std    0.816497    8.539126  2.236068
min    1.000000   10.000000  0.000000
25%    1.750000   13.750000  0.000000
50%    2.000000   17.500000  0.000000
75%    2.250000   22.500000  0.000000
max    3.000000   30.000000  5.000000
Row count and unknown quantities/prices preserved.

The discount count becomes five and its mean changes. Quantity and price counts remain four. Always explain the rule behind a change in statistics.

Common Questions and Mistakes

  • Why is a column missing? The default mixed-data report selects numeric columns. Check dtypes and use include.
  • Why is count smaller than the row count? That column has missing values.
  • Why does top change? Tied most-frequent labels do not have a guaranteed winner.
  • Why is std NaN? There may be fewer than two non-missing observations.
  • Why does selecting text fail? Check whether its dtype is string, object or category. A selection with no matching columns raises an error.
  • Does describe detect outliers? It provides clues; inspect records and use plots before drawing conclusions.

Practice in Google Colab

Run the examples and exercises in the accompanying notebook. Sign in, choose File → Save a copy in Drive, then run the cells in order. The sales CSV loads automatically from Plus2Net.

Open in Google Colab View on GitHub
  1. Describe marks without including ID.
  2. Compare count with the number of rows.
  3. Add missing marks and examine the report.
  4. Compare zero and median filling.
  5. Explain why a mean order ID is not a business metric.
Download sample sales CSV

Summary and Next Steps

Use describe to understand observed values, then investigate missingness and choose justified cleaning rules. Select meaningful measurements, interpret statistics according to dtype, and verify changes after cleaning.

Return to Data Exploration and Analysis, or continue with fillna, DataFrame dtypes, or the Pandas reporting exercises. The exercises connect aggregation, groupby and reporting; describe itself does not merge tables.

Reference: Official Pandas describe documentation.

Pandas Tutorials DataFrame


Subscribe to our YouTube Channel here



plus2net.com







Python Video Tutorials
Python SQLite Video Tutorials
Python MySQL Video Tutorials
Python Tkinter Video Tutorials
✖
We use cookies to improve your browsing experience. . Learn more
HTML MySQL PHP JavaScript ASP Photoshop Articles Contact us
© 2000-2026 plus2net.com All rights reserved worldwide Privacy Policy Disclaimer