portfolios.compute_portfolio_returns()
Compute portfolio returns.
Usage
portfolios.compute_portfolio_returns(
data,
sorting_variables,
sorting_method,
rebalancing_month=None,
breakpoint_options_main=None,
breakpoint_options_secondary=None,
breakpoint_function_main=None,
breakpoint_function_secondary=None,
min_portfolio_size=1,
cap_weight=0.8,
data_options=None,
quiet=False
)Computes individual portfolio returns based on specified sorting variables and sorting methods. The portfolios can be rebalanced every period or on an annual frequency by specifying a rebalancing month, which is only applicable at a monthly return frequency. The function supports univariate and bivariate sorts, with the latter supporting dependent and independent sorting methods.
The function checks for consistency in the provided arguments. For univariate sorts, a single sorting variable and a corresponding number of portfolios must be provided. For bivariate sorts, two sorting variables and two corresponding numbers of portfolios (or percentiles) are required. The sorting method determines how portfolios are assigned and how returns are computed. The function handles missing and extreme values appropriately based on the specified sorting method and rebalancing frequency.
Parameters
data: pl.DataFrame-
Stock-level panel. Must contain the id, date, and excess return columns (configurable via ‘data_options’), the sorting variable(s), and optionally a market-cap lag column for value-weighted returns.
sorting_variables: str or list of str-
Column name(s) in ‘data’ to use for sorting and portfolio assignment. For univariate sorts, provide a single variable. For bivariate sorts, provide two variables, where the first string refers to the main variable and the second string refers to the secondary (‘control’) variable.
sorting_method: {‘univariate’, ‘bivariate-dependent’,-
‘bivariate-independent’} Sorting method to use. For bivariate sorts, the portfolio returns are averaged over the controlling sorting variable (i.e., the second sorting variable), and only portfolio returns for the main sorting variable are returned.
rebalancing_month: int = None-
Integer between 1 and 12 specifying the month in which to form portfolios that are held constant for one year. For example, setting it to 7 creates portfolios in July that are held constant until June of the following year. The default None corresponds to periodic rebalancing.
breakpoint_options_main: dict = None-
Named dict of ‘breakpoint_options’ passed to ‘breakpoint_function_main’ for the main sorting variable. Required.
breakpoint_options_secondary: dict = None-
Named dict of ‘breakpoint_options’ passed to ‘breakpoint_function_secondary’. Required for bivariate sorts.
breakpoint_function_main: callable = None-
Function to compute breakpoints for the main sorting variable. Defaults to ‘compute_breakpoints’.
breakpoint_function_secondary: callable = None-
Function to compute breakpoints for the secondary sorting variable. Defaults to ‘compute_breakpoints’.
min_portfolio_size: int = 1-
Minimum number of firms required in the reported portfolio cross-section on a given date. For univariate sorts that is firms per portfolio-date; for bivariate sorts that is firms per main-portfolio-date summed across the secondary buckets. Cross-sections below the threshold have their returns set to null. A typical value is 5 (the Fama-French convention). Set to 0 to deactivate the check entirely.
cap_weight: float = 0.8-
Quantile of the cross-sectional ‘mktcap_lag’ distribution at which market capitalizations are capped per date when computing the capped value-weighted excess return (‘ret_excess_vw_capped’). Must be in [0, 1].
data_options: dict = None-
Column-name mapping (see ‘data_options’). The ‘id’, ‘date’, ‘ret_excess’, and ‘mktcap_lag’ elements are used. Uses ‘data_options’ defaults when None.
quiet: bool = False- If True, suppress informational warnings about missing observations in the output panel.
Returns
pl.DataFrame-
Data frame with computed portfolio returns as a complete panel (all portfolio-date combinations), containing:
- ‘portfolio’: Portfolio identifier.
- date column (as in ‘data_options’): Date of the portfolio return.
- ‘ret_excess_vw’: Value-weighted excess return (only if ‘data’ contains the market-cap lag column). Null if insufficient observations.
- ‘ret_excess_ew’: Equal-weighted excess return. Null if insufficient observations.
- ‘ret_excess_vw_capped’: Capped value-weighted excess return (only if ‘data’ contains the market-cap lag column). Weights are computed using market capitalization capped at the ‘cap_weight’ percentile per date. Null if insufficient observations.
Notes
Ensure that ‘data’ contains all required columns: the specified sorting variables and excess returns (see options and defaults set in ‘data_options’). A ValueError is raised if any required column is missing.
Examples
import datetime as dt
import numpy as np
import polars as pl
from tidyfinance import (
compute_portfolio_returns,
breakpoint_options,
)
rng = np.random.default_rng(42)
dates = pl.date_range(
dt.date(2020, 1, 1), dt.date(2028, 4, 1), '1mo', eager=True
)
data = pl.DataFrame({
'permno': range(1, 501),
'date': dates.to_numpy().repeat(5),
'mktcap_lag': rng.uniform(100, 1000, 500),
'ret_excess': rng.standard_normal(500),
'size': rng.uniform(50, 150, 500),
})
# Univariate sorting with periodic rebalancing
compute_portfolio_returns(
data, 'size', 'univariate',
breakpoint_options_main=breakpoint_options(n_portfolios=5),
)
# Bivariate dependent sorting with annual rebalancing
compute_portfolio_returns(
data, ['size', 'mktcap_lag'], 'bivariate-dependent', 7,
breakpoint_options_main=breakpoint_options(n_portfolios=5),
breakpoint_options_secondary=breakpoint_options(n_portfolios=3),
)