lagging.join_lagged_values()
Join lagged values of variables over a date range.
Usage
lagging.join_lagged_values(
original_data,
new_data,
id_keys,
min_lag,
max_lag,
ff_adjustment=False,
date_col="date",
data_options=None
)Joins lagged values of selected variables from one data frame (‘new_data’) into another (‘original_data’), based on date ranges defined by ‘min_lag’ and ‘max_lag’. Unlike ‘add_lagged_columns’, this function supports joining across data frames with different date grids (e.g. monthly source data into quarterly target data). All columns in ‘new_data’ besides ‘id_keys’ and the date column are lagged and joined under their original names.
Parameters
original_data: pl.DataFrame-
Target panel data. The date column must be of dtype ‘pl.Date’ or ‘pl.Datetime’.
new_data: pl.DataFrame-
Source variables to lag and merge. All columns besides ‘id_keys’ and the date column will be lagged and joined.
id_keys: list of str or str-
Identifier column(s) shared by both frames.
min_lag: int, str, datetime.timedelta, or pd.DateOffset-
Lower lag bound (inclusive).
max_lag: int, str, datetime.timedelta, or pd.DateOffset-
Upper lag bound (inclusive).
ff_adjustment: bool = False-
If True, keeps only the last observation per identifier and year in ‘new_data’ before lagging (Fama-French convention). Defaults to False.
date_col: str = "date"-
Name of the date column. Defaults to ‘date’.
data_options: dict = None- Column-name mapping (see ‘data_options’). The ‘date’ element is used to identify the date column. Uses the ‘data_options’ default when None: ‘date’ -> ‘date’.
Returns
pl.DataFrame- ‘original_data’ with all columns from ‘new_data’ appended as lagged values (keeping their original names).
Examples
import numpy as np
import polars as pl
import datetime as dt
from tidyfinance import join_lagged_values
rng = np.random.default_rng(42)
dates = pl.date_range(
dt.date(2020, 1, 1), dt.date(2020, 6, 1), "1mo", eager=True
)
df1 = pl.DataFrame({
'id': [1] * 6 + [2] * 6,
'date': dates.to_list() * 2,
})
df2 = df1.with_columns(
x=pl.Series(rng.standard_normal(len(df1)))
)
join_lagged_values(
original_data=df1,
new_data=df2,
id_keys='id',
min_lag='1mo',
max_lag='3mo',
)