portfolios.compute_breakpoints()

Compute breakpoints based on a sorting variable.

Usage

Source

portfolios.compute_breakpoints(
    data, sorting_variable, breakpoint_options, data_options=None
)

Computes breakpoints based on a specified sorting variable. It can optionally filter the data by exchanges or lagged size quantiles before computing the breakpoints. The function requires either the number of portfolios to be created or specific percentiles for the breakpoints, but not both. The function also optionally handles cases where the sorting variable clusters on the edges, by assigning all extreme values to the edges and attempting to compute equally populated breakpoints with the remaining values.

Parameters

data: pl.DataFrame

DataFrame with the dataset for breakpoint computation.

sorting_variable: str

Column name in ‘data’ to be used for determining breakpoints.

breakpoint_options: dict

Named dict of ‘breakpoint_options’ for the breakpoints. The accepted entries include:

  • ‘n_portfolios’ (int, optional): Number of equally sized portfolios to create. Mutually exclusive with ‘percentiles’.
  • ‘percentiles’ (list of float, optional): Percentiles defining the breakpoints of the portfolios. Mutually exclusive with ‘n_portfolios’.
  • ‘breakpoints_exchanges’ (str or list of str, optional): Exchange names to filter the data before computing breakpoints. Exchanges must be stored in the column given by ‘data_options’ (defaults to ‘exchange’). If None, no filtering is applied.
  • ‘smooth_bunching’ (bool, optional): Whether to attempt smoothing non-extreme portfolios if the sorting variable bunches on the extremes (True) or not (False, the default). In some cases, smoothing will not result in equal-sized portfolios off the edges due to multiple clusters. If sufficiently large bunching is detected, ‘percentiles’ is ignored and equally-spaced portfolios are returned for these cases with a warning.
  • ‘breakpoints_min_size_threshold’ (float, optional): Value between 0 and 1 (exclusive). When set, stocks with market capitalization below this quantile are excluded from breakpoint computation. The quantile is computed among ‘breakpoints_exchanges’ stocks if specified, otherwise among all stocks. Requires a market capitalization column in the data (column name determined by ‘data_options’).
data_options: dict = None
Column-name mapping (see ‘data_options’). The ‘exchange’ key is used to specify the exchange column, and ‘mktcap_lag’ is used to specify the market capitalization column. Uses the ‘data_options’ default when None: ‘exchange’ -> ‘exchange’ and ‘mktcap_lag’ -> ‘mktcap_lag’.

Returns

np.ndarray
Sorted array of breakpoints of the desired length.

Notes

This function raises a ValueError if both ‘n_portfolios’ and ‘percentiles’ are provided or missing simultaneously.

Examples

import numpy as np
import polars as pl
from tidyfinance import compute_breakpoints, breakpoint_options
rng = np.random.default_rng(42)
data = pl.DataFrame({
    'id': range(1, 101),
    'exchange': rng.choice(['NYSE', 'NASDAQ'], 100),
    'market_cap': range(1, 101),
})
compute_breakpoints(
    data, 'market_cap', breakpoint_options(n_portfolios=5)
)
compute_breakpoints(
    data,
    'market_cap',
    breakpoint_options(
        percentiles=[0.2, 0.4, 0.6, 0.8],
        breakpoints_exchanges=['NYSE'],
    ),
)