utilities.create_summary_statistics()

Create summary statistics for specified variables.

Usage

Source

utilities.create_summary_statistics(
    data, variables, by=None, detail=False, drop_na=False
)

Computes a set of summary statistics for numeric and boolean variables in a data frame. It allows users to select specific variables for summarization and can calculate statistics for the whole dataset or within groups specified by the ‘by’ argument. Additional detail levels for quantiles can be included.

The function first checks that all specified variables are of a numeric dtype (int, float, or bool). If any variables fail this check, a ‘ValueError’ is raised listing the offending columns. Boolean columns are summarized as their numeric equivalent — for example, the ‘mean’ of a boolean column is the proportion of True.

The basic set of summary statistics includes the count of non-null values (‘n’), ‘mean’, standard deviation (‘sd’), minimum (‘min’), median (‘q50’), and maximum (‘max’). If ‘detail’ is True, the function also computes the 1st, 5th, 10th, 25th, 75th, 90th, 95th, and 99th percentiles (‘q01’ through ‘q99’). Statistics are computed for the whole dataset, or separately for each group when ‘by’ is supplied.

Parameters

data: pl.DataFrame

Data frame containing the variables to be summarized.

variables: list of str

List of column names in the data frame to summarize. These variables must be of a numeric dtype (int, float, or bool).

by: str = None

Column name to group the data before summarizing. If None (the default), summary statistics are computed across all observations.

detail: bool = False

Whether to compute detailed summary statistics, including additional quantiles. When False, computes basic statistics (n, mean, sd, min, q50, max). When True, additional quantiles (q01, q05, q10, q25, q75, q90, q95, q99) are computed.

drop_na: bool = False
Whether to drop missing values for each variable before summarizing.

Returns

pl.DataFrame
Data frame with summary statistics for each selected variable. If ‘by’ is specified, the output includes the grouping variable as well. Each row represents a variable (and a group if ‘by’ is used), and each column contains the computed statistics.

Examples

import polars as pl
from tidyfinance import create_summary_statistics
data = pl.DataFrame({
    'ret': [0.01, -0.02, 0.03, None, 0.005],
    'size': [100, 200, 150, 300, 250],
    'group': ['A', 'A', 'B', 'B', 'A'],
})
# Basic summary across all observations
create_summary_statistics(data, ['ret', 'size'])
# Grouped summary
create_summary_statistics(data, ['ret', 'size'], by='group')
# Detailed quantiles
create_summary_statistics(data, ['ret'], detail=True)