regression.estimate_fama_macbeth()
Estimate Fama-MacBeth regressions.
Usage
regression.estimate_fama_macbeth(
data,
model,
vcov="newey-west",
vcov_options=None,
date_col="date",
data_options=None,
detail=False
)Runs one cross-sectional ordinary least squares regression per period of ‘date_col’, then averages the per-period coefficients to obtain risk premia and aggregates them into a single tidy frame.
Parameters
data: pl.DataFrame-
Panel containing the dependent and independent variables named in ‘model’ plus a column with the time index. Each (date, unit) combination should appear at most once.
model: str-
Formula describing the cross-sectional regression (e.g., ‘ret_excess ~ beta + bm + log_mktcap’). Standard formulaic syntax; an intercept is included unless the formula ends in ‘- 1’.
vcov: (iid, newey - west) = "iid"-
Standard error treatment for the time-series average of period coefficients. ‘iid’ assumes independent and identically distributed errors across periods. ‘newey-west’ applies Newey-West heteroskedasticity- and autocorrelation-consistent standard errors with Bartlett kernel.
vcov_options: dict = None-
Tuning options for the Newey-West estimator. Recognized keys:
- ‘lag’ : int, optional Bartlett truncation lag. If None (the default), the automatic bandwidth from Newey & West (1994) is used.
- ‘prewhite’ : int, default 1 Order of the VAR prewhitening filter applied before computing the long-run variance. Pass 0 to disable.
- ‘adjust’ : bool, default False Apply a finite-sample degrees-of-freedom correction.
- ‘maxlags’ : int, optional Deprecated alias for ‘lag’ (with ‘prewhite’ defaulting to 0). Emits a DeprecationWarning.
date_col: str = "date"-
Column in ‘data’ identifying the time index for cross-sectional regressions.
data_options: dict = None-
Column-name mapping (see ‘data_options’). The ‘date’ element is used to specify the date column, overriding ‘date_col’. Uses the ‘data_options’ default when None: ‘date’ -> ‘date’.
detail: bool = False- If ‘False’ (default), return only the coefficient estimates. If ‘True’, return a dict with two keys: ‘coefficients’ (the usual estimates data frame) and ‘summary_statistics’ (a one-row data frame with the average cross-sectional R-squared, adjusted R-squared, and number of observations per cross-section).
Returns
pl.DataFrame or dict-
If ‘detail’ is ‘False’ (default), a data frame with one row per term in ‘model’, in model-term order (the intercept first, then the regressors as they appear in the formula), with columns:
- ‘factor’ : term name (‘intercept’ or a regressor)
- ‘risk_premium’ : time-series mean of cross-sectional coefficients
- ‘n’ : number of periods used
- ‘standard_error’ : SE of the time-series mean under ‘vcov’
- ‘t_statistic’ : risk_premium / standard_error
If ‘detail’ is ‘True’, a dict with two elements:
- ‘coefficients’ : the same data frame described above
- ‘summary_statistics’ : a one-row data frame with ‘r_squared’ (mean cross-sectional R-squared), ‘adj_r_squared’ (mean cross-sectional adjusted R-squared), and ‘n_obs’ (mean cross-sectional observation count)
Raises
ValueError- If ‘vcov’ is not ‘iid’ or ‘newey-west’, if ‘vcov_options’ contains an unrecognized key, if ‘date_col’ is missing from ‘data’, or if any date grouping has too few rows to estimate the cross-sectional coefficients (each grouping needs more rows than the number of variables in ‘model’).
References
Fama, E. F., and MacBeth, J. D. (1973). Risk, return, and equilibrium: Empirical tests. Journal of Political Economy, 81(3), 607-636. https://doi.org/10.1086/260061
Newey, W. K., and West, K. D. (1987). A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix. Econometrica, 55(3), 703-708. https://doi.org/10.2307/1913610
Newey, W. K., and West, K. D. (1994). Automatic lag selection in covariance matrix estimation. Review of Economic Studies, 61(4), 631-653. https://doi.org/10.2307/2297912
Examples
from datetime import date
import numpy as np
import polars as pl
from tidyfinance import estimate_fama_macbeth
rng = np.random.default_rng(1234)
dates = pl.date_range(
date(2020, 1, 1), date(2020, 12, 1), "1mo", eager=True
)
data = pl.DataFrame({
'date': np.repeat(dates.to_numpy(), 50),
'permno': np.tile(np.arange(1, 51), 12),
'ret_excess': rng.normal(0, 0.1, 600),
'beta': rng.normal(1, 0.2, 600),
'bm': rng.normal(0.5, 0.1, 600),
'log_mktcap': rng.normal(10, 1, 600),
})
result = estimate_fama_macbeth(data, 'ret_excess ~ beta+bm+log_mktcap')
# Override the Newey-West settings
result_iid = estimate_fama_macbeth(
data,
'ret_excess ~ beta + bm + log_mktcap',
vcov='iid',
)
# Return detailed output including R-squared and observation counts
result_detail = estimate_fama_macbeth(
data,
'ret_excess ~ beta + bm + log_mktcap',
detail=True,
)