utilities.create_summary_statistics()
Create summary statistics for specified variables.
Usage
utilities.create_summary_statistics(
data, variables, by=None, detail=False, drop_na=False
)Computes a set of summary statistics for numeric and boolean variables in a data frame. It allows users to select specific variables for summarization and can calculate statistics for the whole dataset or within groups specified by the ‘by’ argument. Additional detail levels for quantiles can be included.
The function first checks that all specified variables are of a numeric dtype (int, float, or bool). If any variables fail this check, a ‘ValueError’ is raised listing the offending columns. Boolean columns are summarized as their numeric equivalent — for example, the ‘mean’ of a boolean column is the proportion of True.
The basic set of summary statistics includes the count of non-null values (‘n’), ‘mean’, standard deviation (‘sd’), minimum (‘min’), median (‘q50’), and maximum (‘max’). If ‘detail’ is True, the function also computes the 1st, 5th, 10th, 25th, 75th, 90th, 95th, and 99th percentiles (‘q01’ through ‘q99’). Statistics are computed for the whole dataset, or separately for each group when ‘by’ is supplied.
Parameters
data: pl.DataFrame-
Data frame containing the variables to be summarized.
variables: list of str-
List of column names in the data frame to summarize. These variables must be of a numeric dtype (int, float, or bool).
by: str = None-
Column name to group the data before summarizing. If None (the default), summary statistics are computed across all observations.
detail: bool = False-
Whether to compute detailed summary statistics, including additional quantiles. When False, computes basic statistics (n, mean, sd, min, q50, max). When True, additional quantiles (q01, q05, q10, q25, q75, q90, q95, q99) are computed.
drop_na: bool = False- Whether to drop missing values for each variable before summarizing.
Returns
pl.DataFrame- Data frame with summary statistics for each selected variable. If ‘by’ is specified, the output includes the grouping variable as well. Each row represents a variable (and a group if ‘by’ is used), and each column contains the computed statistics.
Examples
import polars as pl
from tidyfinance import create_summary_statistics
data = pl.DataFrame({
'ret': [0.01, -0.02, 0.03, None, 0.005],
'size': [100, 200, 150, 300, 250],
'group': ['A', 'A', 'B', 'B', 'A'],
})
# Basic summary across all observations
create_summary_statistics(data, ['ret', 'size'])
# Grouped summary
create_summary_statistics(data, ['ret', 'size'], by='group')
# Detailed quantiles
create_summary_statistics(data, ['ret'], detail=True)