regression.estimate_fama_macbeth()

Estimate Fama-MacBeth regressions.

Usage

Source

regression.estimate_fama_macbeth(
    data,
    model,
    vcov="newey-west",
    vcov_options=None,
    date_col="date",
    data_options=None,
    detail=False
)

Runs one cross-sectional ordinary least squares regression per period of ‘date_col’, then averages the per-period coefficients to obtain risk premia and aggregates them into a single tidy frame.

Parameters

data: pl.DataFrame

Panel containing the dependent and independent variables named in ‘model’ plus a column with the time index. Each (date, unit) combination should appear at most once.

model: str

Formula describing the cross-sectional regression (e.g., ‘ret_excess ~ beta + bm + log_mktcap’). Standard formulaic syntax; an intercept is included unless the formula ends in ‘- 1’.

vcov: (iid, newey - west) = "iid"

Standard error treatment for the time-series average of period coefficients. ‘iid’ assumes independent and identically distributed errors across periods. ‘newey-west’ applies Newey-West heteroskedasticity- and autocorrelation-consistent standard errors with Bartlett kernel.

vcov_options: dict = None

Tuning options for the Newey-West estimator. Recognized keys:

  • ‘lag’ : int, optional Bartlett truncation lag. If None (the default), the automatic bandwidth from Newey & West (1994) is used.
  • ‘prewhite’ : int, default 1 Order of the VAR prewhitening filter applied before computing the long-run variance. Pass 0 to disable.
  • ‘adjust’ : bool, default False Apply a finite-sample degrees-of-freedom correction.
  • ‘maxlags’ : int, optional Deprecated alias for ‘lag’ (with ‘prewhite’ defaulting to 0). Emits a DeprecationWarning.
date_col: str = "date"

Column in ‘data’ identifying the time index for cross-sectional regressions.

data_options: dict = None

Column-name mapping (see ‘data_options’). The ‘date’ element is used to specify the date column, overriding ‘date_col’. Uses the ‘data_options’ default when None: ‘date’ -> ‘date’.

detail: bool = False
If ‘False’ (default), return only the coefficient estimates. If ‘True’, return a dict with two keys: ‘coefficients’ (the usual estimates data frame) and ‘summary_statistics’ (a one-row data frame with the average cross-sectional R-squared, adjusted R-squared, and number of observations per cross-section).

Returns

pl.DataFrame or dict

If ‘detail’ is ‘False’ (default), a data frame with one row per term in ‘model’, in model-term order (the intercept first, then the regressors as they appear in the formula), with columns:

  • ‘factor’ : term name (‘intercept’ or a regressor)
  • ‘risk_premium’ : time-series mean of cross-sectional coefficients
  • ‘n’ : number of periods used
  • ‘standard_error’ : SE of the time-series mean under ‘vcov’
  • ‘t_statistic’ : risk_premium / standard_error

If ‘detail’ is ‘True’, a dict with two elements:

  • ‘coefficients’ : the same data frame described above
  • ‘summary_statistics’ : a one-row data frame with ‘r_squared’ (mean cross-sectional R-squared), ‘adj_r_squared’ (mean cross-sectional adjusted R-squared), and ‘n_obs’ (mean cross-sectional observation count)

Raises

ValueError
If ‘vcov’ is not ‘iid’ or ‘newey-west’, if ‘vcov_options’ contains an unrecognized key, if ‘date_col’ is missing from ‘data’, or if any date grouping has too few rows to estimate the cross-sectional coefficients (each grouping needs more rows than the number of variables in ‘model’).

References

Fama, E. F., and MacBeth, J. D. (1973). Risk, return, and equilibrium: Empirical tests. Journal of Political Economy, 81(3), 607-636. https://doi.org/10.1086/260061

Newey, W. K., and West, K. D. (1987). A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix. Econometrica, 55(3), 703-708. https://doi.org/10.2307/1913610

Newey, W. K., and West, K. D. (1994). Automatic lag selection in covariance matrix estimation. Review of Economic Studies, 61(4), 631-653. https://doi.org/10.2307/2297912

Examples

from datetime import date
import numpy as np
import polars as pl
from tidyfinance import estimate_fama_macbeth
rng = np.random.default_rng(1234)
dates = pl.date_range(
    date(2020, 1, 1), date(2020, 12, 1), "1mo", eager=True
)
data = pl.DataFrame({
    'date': np.repeat(dates.to_numpy(), 50),
    'permno': np.tile(np.arange(1, 51), 12),
    'ret_excess': rng.normal(0, 0.1, 600),
    'beta': rng.normal(1, 0.2, 600),
    'bm': rng.normal(0.5, 0.1, 600),
    'log_mktcap': rng.normal(10, 1, 600),
})
result = estimate_fama_macbeth(data, 'ret_excess ~ beta+bm+log_mktcap')
# Override the Newey-West settings
result_iid = estimate_fama_macbeth(
    data,
    'ret_excess ~ beta + bm + log_mktcap',
    vcov='iid',
)
# Return detailed output including R-squared and observation counts
result_detail = estimate_fama_macbeth(
    data,
    'ret_excess ~ beta + bm + log_mktcap',
    detail=True,
)