portfolios.compute_breakpoints()
Compute breakpoints based on a sorting variable.
Usage
portfolios.compute_breakpoints(
data, sorting_variable, breakpoint_options, data_options=None
)Computes breakpoints based on a specified sorting variable. It can optionally filter the data by exchanges or lagged size quantiles before computing the breakpoints. The function requires either the number of portfolios to be created or specific percentiles for the breakpoints, but not both. The function also optionally handles cases where the sorting variable clusters on the edges, by assigning all extreme values to the edges and attempting to compute equally populated breakpoints with the remaining values.
Parameters
data: pl.DataFrame-
DataFrame with the dataset for breakpoint computation.
sorting_variable: str-
Column name in ‘data’ to be used for determining breakpoints.
breakpoint_options: dict-
Named dict of ‘breakpoint_options’ for the breakpoints. The accepted entries include:
- ‘n_portfolios’ (int, optional): Number of equally sized portfolios to create. Mutually exclusive with ‘percentiles’.
- ‘percentiles’ (list of float, optional): Percentiles defining the breakpoints of the portfolios. Mutually exclusive with ‘n_portfolios’.
- ‘breakpoints_exchanges’ (str or list of str, optional): Exchange names to filter the data before computing breakpoints. Exchanges must be stored in the column given by ‘data_options’ (defaults to ‘exchange’). If None, no filtering is applied.
- ‘smooth_bunching’ (bool, optional): Whether to attempt smoothing non-extreme portfolios if the sorting variable bunches on the extremes (True) or not (False, the default). In some cases, smoothing will not result in equal-sized portfolios off the edges due to multiple clusters. If sufficiently large bunching is detected, ‘percentiles’ is ignored and equally-spaced portfolios are returned for these cases with a warning.
- ‘breakpoints_min_size_threshold’ (float, optional): Value between 0 and 1 (exclusive). When set, stocks with market capitalization below this quantile are excluded from breakpoint computation. The quantile is computed among ‘breakpoints_exchanges’ stocks if specified, otherwise among all stocks. Requires a market capitalization column in the data (column name determined by ‘data_options’).
data_options: dict = None- Column-name mapping (see ‘data_options’). The ‘exchange’ key is used to specify the exchange column, and ‘mktcap_lag’ is used to specify the market capitalization column. Uses the ‘data_options’ default when None: ‘exchange’ -> ‘exchange’ and ‘mktcap_lag’ -> ‘mktcap_lag’.
Returns
np.ndarray- Sorted array of breakpoints of the desired length.
Notes
This function raises a ValueError if both ‘n_portfolios’ and ‘percentiles’ are provided or missing simultaneously.
Examples
import numpy as np
import polars as pl
from tidyfinance import compute_breakpoints, breakpoint_options
rng = np.random.default_rng(42)
data = pl.DataFrame({
'id': range(1, 101),
'exchange': rng.choice(['NYSE', 'NASDAQ'], 100),
'market_cap': range(1, 101),
})
compute_breakpoints(
data, 'market_cap', breakpoint_options(n_portfolios=5)
)
compute_breakpoints(
data,
'market_cap',
breakpoint_options(
percentiles=[0.2, 0.4, 0.6, 0.8],
breakpoints_exchanges=['NYSE'],
),
)