download.download_data()
Download and process data based on domain and dataset.
Usage
download.download_data(
domain=None,
dataset=None,
start_date=None,
end_date=None,
type=None,
**kwargs
)Downloads and processes data based on the specified domain (e.g., Fama-French factors, Global Q factors, or macro predictors), dataset, and date range. The function checks whether the specified domain is supported and then delegates to the appropriate function for downloading and processing the data.
Parameters
domain: str = None-
The domain of the dataset to download, given as one of the canonical names returned by ‘list_supported_datasets()’: ‘Fama-French’, ‘Global Q’, ‘Goyal-Welch’, ‘WRDS’, ‘Pseudo Data’, ‘Index Constituents’, ‘FRED’, ‘Stock Prices’, ‘Open Source Asset Pricing’, ‘Global Factor Data’, ‘Pastor-Stambaugh’, ‘Stambaugh-Yuan’, ‘Tidy Finance’. The previous short names (e.g. ‘famafrench’, ‘wrds’, ‘pseudo’) are still accepted but deprecated and will be removed in a future release.
dataset: str = None-
The specific dataset to download within the domain.
start_date: str = None-
A character string or date in ‘YYYY-MM-DD’ format specifying the start date for the data. If not provided, the full dataset or a subset is returned, depending on the dataset type.
end_date: str = None-
A character string or date in ‘YYYY-MM-DD’ format specifying the end date for the data. If not provided, the full dataset or a subset is returned, depending on the dataset type.
type: str = None-
Deprecated. Use ‘domain’ and ‘dataset’ instead. If provided, a DeprecationWarning is emitted and the legacy type is translated to a (‘domain’, ‘dataset’) pair via ‘list_supported_datasets’.
**kwargs- Additional arguments passed to specific download functions depending on ‘domain’. For instance, if ‘domain’ is ‘Index Constituents’, arguments are passed to ’_download_data_constituents’. If ‘domain’ is ‘Global Factor Data’, the ‘dataset’ argument and arguments such as ‘region’, ‘factors’, ‘classification’, ‘frequency’, and ‘weighting’ are passed to ’_download_data_jkp’. If ‘domain’ is ‘Tidy Finance’ and ‘dataset’ is ‘factor_library’, arguments are either filter inputs (e.g., ‘sorting_variable’, ‘rebalancing’, ‘fill_all’) or an explicit ‘ids’ vector that bypasses the grid filter and downloads the specified portfolios directly via ’_download_factor_library_ids’; see ’_download_data_huggingface’ for details.
Returns
pl.DataFrame- A data frame with processed data, including dates and the relevant financial metrics, filtered by the specified date range. For ‘Stock Prices’, dates are trading days in the exchange’s local time zone (UTC when Yahoo Finance omits it), and both start_date and end_date are inclusive.
Examples
from tidyfinance import download_data
download_data(
'Fama-French',
'Fama/French 5 Factors (2x3) [Daily]',
'2000-01-01',
'2020-12-31',
)
download_data(
'Goyal-Welch', 'monthly', '2000-01-01', '2020-12-31'
)
download_data('Index Constituents', index='DAX')
download_data('FRED', series=['GDP', 'CPIAUCNS'])
download_data('Stock Prices', symbols=['AAPL', 'MSFT'])
download_data(
'Global Factor Data',
region='usa',
factors='mkt',
start_date='2000-01-01',
end_date='2020-12-31',
)
download_data(
'Pastor-Stambaugh', '2020-01-01', '2020-12-31'
)
download_data(
'Stambaugh-Yuan', 'monthly', '2015-01-01', '2016-12-31'
)
download_data(
'Tidy Finance', 'risk_free', '2020-01-01', '2020-12-31'
)
download_data(
'Tidy Finance',
'high_frequency_sp500',
'2007-07-26',
'2007-07-27',
)
download_data(
'Tidy Finance',
'factor_library',
sorting_variable='high52',
rebalancing='annual',
)
download_data('Tidy Finance', 'factor_library', ids=[1, 2, 3])
download_data('Tidy Finance', 'factor_library_grid')