Changelog
This changelog is generated automatically from GitHub Releases.
v0.5.2
2026-10-01 · GitHub
v0.5.2 (2026-10-01)
- Fixed (Yahoo Finance dates): Stock price downloads now use the exchange’s local time zone, falling back to UTC when the metadata is missing or empty. This corrects previous-day dates for markets such as Australia and New Zealand. Requests include a two-day buffer and results are filtered to include both
start_dateandend_date, matching r-tidyfinance [#305](https://github.com/tidy-finance/py-tidyfinance/issues/305). An end date of today can include the current, still-forming daily bar. - Fixed (factor library layout):
download_data("Tidy Finance", "factor_library", ...)reads the new layout of the factor library on Hugging Face, where the returns hold onlyid,date, andretin files of 1,000 consecutive IDs named after their range (e.g.,id_0000001-0001000.parquet). The file that holds each requested ID is computed from the ID, so only those files are downloaded, each once. Neither dataset is listed any more: the grid is read by name fromportfolio_sort_grid.parquet, so it keeps working when the grid dataset gains more files. The result has the columnsid,date, andretfollowed by the grid columns; it no longer has aret_typecolumn, and the weighting scheme is in theweighting_schemecolumn of the grid. Returns are stored in single precision and returned as 64-bit floats. IDs without returns, whose portfolio sort produced no portfolios, are absent from the result with a warning. Version 0.5.1 cannot read the new layout. This follows r-tidyfinance [#306](https://github.com/tidy-finance/py-tidyfinance/issues/306). - Breaking (factor library sorting variables): The factor library now builds on the signals of Open Source Asset Pricing, so sorting variables carry their names, e.g.
"size"instead of"me","high52"instead of"52w", and"assetgrowth"instead of"ag". The grid covers 179 sorting variables and adds the"1m"lag, the"bivariate-dependent"and"bivariate-independent"sorting methods, and"capped VW"weighting. The examples and documented grid values follow the new release.
v0.5.1
2026-09-01 · GitHub
v0.5.1 (2026-09-01)
- Fixed: Goyal-Welch (macro predictor) CSV parse no longer fails when the S&P Index column uses thousands separators (e.g.
"1,049.34"). Rows after Polars’ default schema-inference window were previously dropped as an empty download.
v0.5.0
2026-08-05 · GitHub
v0.5.0 (2026-07-29)
- Polars rewrite: the entire package is now implemented in polars. All module-level functions take and return
polars.DataFrameobjects, calendar-date columns are typed aspolars.Date(matching the R package and the book), and thepolarsbackend (set_backend("polars"), as used throughout the Tidy Finance book) is a zero-overhead pass-through. The public API and the defaultpandasbackend are unchanged: pandas users keep receiving pandas frames, with conversion now happening at the package boundary (pandas in → polars internals → pandas out).polarsmoved from an optional extra to a core dependency. - Missing values are nulls: data frame outputs now represent missing values as polars nulls (surfacing as
NaN/NaTafter conversion under the pandas backend), instead of floatNaNsentinels. - Breaking (create_summary_statistics): output columns now follow the R package naming —
n,mean,sd,min,q50,max(plusq01…q99withdetail=True) — and grouped summaries return a tidy long table (one row per group × variable) instead of pandas MultiIndex columns. - Breaking (assign_portfolio): returns a
polars.Seriesunder the polars backend (apandas.Seriesunder the default pandas backend, as before). - Lag arguments accept polars duration strings: add_lagged_columns / join_lagged_values accept lags as polars offsets (e.g.
"1mo","1y2mo") in addition to ints (days),datetime.timedelta/pd.Timedelta, and calendarpd.DateOffsetobjects. - Internal SQL and HTTP readers now parse directly into polars (
polars.read_databasefor WRDS,polars.read_csv/read_parqueton fetched bytes elsewhere). - Dependencies trimmed: dropped
requests(unused since v0.2.2 — all HTTP goes throughcurl_cffi); replaced thedotenvwrapper package with its actual implementationpython-dotenv; movedlxmlto the new optionalscrapingextra (it only powers get_available_famafrench_datasets, which now raises a clear ImportError pointing atpip install tidyfinance[scraping]when lxml is absent). Thepolarsextra is kept as a no-op for backward compatibility. - Breaking (estimate_betas), aligned with r-tidyfinance: coefficient columns are now named
interceptandbeta_<variable>(wasInterceptand the bare variable name), and the identifier and date columns come first, so the output is<id>, date, intercept, beta_<variable>, .... - Deprecated (estimate_betas
lookback), aligned with r-tidyfinance:lookbacknow accepts a duration string —"60mo","30d", and the sub-day units"h","m","s"— which rolls over calendar periods exactly as R’smonths(60)does. Observations falling in the same period are pooled, so the result has one row per identifier and period withdatefloored to the period start; for daily data a"3mo"window therefore yields a monthly beta series fitted on every observation in the trailing three months. Passing a plain integer keeps the previous positional window (a count of consecutive observations, one output row per input row) and now emits aDeprecationWarning. The two agree on a gap-free panel with one observation per period, so monthly examples are unaffected; they differ when a panel has holes. - Fix (estimate_betas default
min_obs): the default is nowround(0.8 * lookback)as in r-tidyfinance, rather than truncating. Forlookback6 this gives 5 instead of 4;lookback60 is unchanged at 48. - Breaking (estimate_fama_macbeth), aligned with r-tidyfinance: the column order is now
factor, risk_premium, n, standard_error, t_statistic(nmoved from last to third), the intercept row is labelledintercept(wasIntercept, and now matches estimate_model), and rows are returned in model-term order — the intercept first, then the regressors as they appear in the formula — instead of alphabetically. - Breaking (estimate_betas): duration lookbacks drop windows below
min_obsinstead of returning null rows. - Breaking (estimate_fama_macbeth): too-small date groups and unknown
vcov_optionskeys raise; added data_options. - Breaking (compute_rolling_value): callbacks receive the active backend’s frame type.
- Breaking (TRACE): pre-2012 messages restricted to
trc_st = 'T'. - Breaking (FF breakpoints): column names are strings (
"0-5"), not tuples. - Fixed:
pd.DateOffset(n)day lags; Series index alignment (pandas backend); CRSP v1 ccm-link join; one-sidedrisk_freedate ranges;NaNas missing in filter_sorting_data; WRDSnumericcolumns cast to float.
v0.4.0
2026-07-27 · GitHub
- Polars backend returns WRDS date columns as
Date(#66): WRDS calendar-date columns (datadate,trd_exctn_dt, CCM link dates, FISD dates) are now cast topolars.Dateinstead ofpolars.Datetime, matching the R package. - Added FRED-MD and FRED-QD macroeconomic databases:
download_data("FRED", "FRED-MD" / "FRED-QD")download the McCracken and Ng (2016, 2021) macro panels.transform=Trueapplies each series’ transform code;vintageenables point-in-time analysis. - Added Global Factor Data, Pastor-Stambaugh, and Stambaugh-Yuan downloads:
download_data("Global Factor Data")downloads portfolios, industries, or cutoffs from Jensen, Kelly, and Pedersen (2023).download_data("Pastor-Stambaugh")anddownload_data("Stambaugh-Yuan")download the liquidity and mispricing factors. - OSAP download aligned with beginning-of-month and scaled returns:
download_data("Open Source Asset Pricing")now uses beginning-of-month dates (was end-of-month) and decimal returns (divided by 100). sorting_variableis now optional forfactor_library: it returns the default construction for all sorting variables when omitted, and passingNonefor a filter column removes that filter.- Added
detailparameter to estimate_fama_macbeth:detail=Truereturns coefficients plus summary statistics (meanr_squared,adj_r_squared,n_obs). Default unchanged. - Dependencies (replaced pyfixest with formulaic): dropped
pyfixestforformulaic.
v0.3.0
2026-06-27 · GitHub
- Dependencies (removed statsmodels): The
statsmodelsdependency was dropped; the regression-based functions now usepyfixestinstead. estimate_model and the cross-sectional / IID-variance steps of estimate_fama_macbeth callpyfixest.feols, and estimate_betas was rewritten to estimate rolling betas via closed-form OLS on cumulative cross-product sums (the design Gram matrixX'Xand moment vectorX'yare accumulated and rolled by cumulative-sum differencing, then solved once per window). This follows the fast beta estimation approach, generalized to multiple regressors, and returns coefficients identical to ordinary least squares while avoiding a full refit per window (#49). - Docs (Great Docs): Added a Great Docs documentation site configured via
great-docs.yml, including LLM-friendly artifacts (llms.txt,llms-full.txt). The API reference is generated from the numpydoc docstrings; build locally withgreat-docs build(on Windows setPYTHONUTF8=1to avoid a cp1252 decode error during post-processing). The generatedgreat-docs/build directory is gitignored. (#29) - Breaking (Python version): The minimum supported Python is now 3.11 (was 3.10), as required by the Great Docs toolchain.
- Docs (R parity): Fixed docstring discrepancies surfaced by the rendered reference, aligning the Python docs with r-tidyfinance: breakpoint_options (removed a duplicated
breakpoints_exchangesentry and documented the previously undocumentedbreakpoints_min_size_threshold), create_summary_statistics (enumerated the reported statistics and detail quantiles), compute_portfolio_returns / implement_portfolio_sort (min_portfolio_sizeunivariate/bivariate semantics and the “set to 0 to deactivate” behavior), estimate_betas (lookbackannotated asintto match its use as an observation-count window), and winsorize (corrected thextype tonp.ndarrayand documented the[0, 0.5]range forcut). - Polars support: the public API can now work with polars data frames via a global backend. Call
tidyfinance.set_backend("polars")(default"pandas"; get_backend() reports the current setting). When set to"polars", the data-bearing functions (download_data, theestimate_*/compute_*family, add_lagged_columns, assign_portfolio’s frame inputs, list_supported_datasets, etc.) return polars data frames, and all of them also accept polars input regardless of the active backend (converted to pandas internally). DataFrame outputs convert; Series/dict/ndarray returns (e.g. assign_portfolio) are left as-is, and date indices are preserved as columns. Internals remain pandas-based for now. Requires the optionalpolarsdependency (pip install tidyfinance[polars]) (#42). - Breaking (WRDS credentials): WRDS credentials are now read exclusively from environment variables (e.g. via a
.envfile). Support forconfig.yamlhas been removed: set_wrds_credentials() now writes a.envfile (withWRDS_USERandWRDS_PASSWORD), and get_wrds_connection() no longer accepts aconfig_pathargument. Thepyyamldependency was dropped. Migrate any existingconfig.yamlcredentials into a.envfile or environment variables. - Breaking (CRSP): the monthly CRSP price column returned by
download_data(domain="wrds", dataset="crsp_monthly")is now namedprc(wasaltprc), aligning with r-tidyfinance and both book editions. The value is unchanged — it ismthprcfrom the CRSP v2 monthly stock file;altprcwas the legacy (v1) column name and was semantically stale for v2 downloads. Update any downstream code that referencedaltprc(including the dependentmktcapcomputation). - Fix (Fama-MacBeth Newey-West): estimate_fama_macbeth now matches R’s
sandwich::NeweyWestdefaults, so the Python and R editions agree on Newey-West t-statistics. The previous implementation used statsmodels HAC with a fixedmaxlags=6and no prewhitening (textbook Newey-West 1987); the new numpy implementation uses VAR(1) prewhitening plus the automatic Newey & West (1994) bandwidth, Bartlett kernel, recoloring, and no finite-sample adjustment (verified againstsandwich3.1.1 to ~1e-13).vcov_optionsnow mirrors R’s interface (lag,prewhite,adjust) and defaults toNone; the legacymaxlagskey is accepted as a deprecated alias forlag(preserving the old no-prewhitening behavior) and emits aDeprecationWarning(#35). - Fix (CRSP column order):
download_data(domain="wrds", dataset="crsp_monthly")now orderslisting_agebeforemktcapto match r-tidyfinance’sdownload_data_wrds_crsp()(..., siccd, listing_age, mktcap, mktcap_lag, ...). Values are unchanged; only the column order differed (#36). - Fix (TRACE regime cutoff): process_trace_data now uses the correct Dick-Nielsen (2014) enhanced-TRACE regime cutoff of
2012-02-06(was the transposed2012-06-02). Samples spanning Feb 6 – Jun 2, 2012 were previously cleaned under the wrong cancellation/correction/reversal regime, producing incorrect output; samples entirely after June 2012 were unaffected. This aligns the Python edition with r-tidyfinance’sdownload_data_wrds_trace_enhanced()(#34). - download_data() now uses the human-readable domain names returned by list_supported_datasets() (e.g.,
"Fama-French","Global Q","WRDS","Tidy Finance"). The"pseudo"and"tidyfinance"domains were renamed to"Pseudo Data"and"Tidy Finance". The previous machine-readable domain names (e.g.,"famafrench","wrds","pseudo","tidyfinance") are soft-deprecated but still accepted. - Breaking (package API): the dataset-specific
_download_data_*helpers (e.g._download_data_wrds,_download_data_macro_predictors,_download_data_constituents,_download_data_factors_ff,_download_data_factors_q,_download_data_osap,_download_data_risk_free,_download_data_stock_prices) are no longer re-exported from the package root. Public access continues via the dispatcherdownload_data(domain, dataset, ...). If you need a helper directly, import it from its defining module (e.g.from tidyfinance.data_download import _download_data_wrds).
Version 0.1.1
2025-03-21 · GitHub
Initial PyPI release