lagging.join_lagged_values()

Join lagged values of variables over a date range.

Usage

Source

lagging.join_lagged_values(
    original_data,
    new_data,
    id_keys,
    min_lag,
    max_lag,
    ff_adjustment=False,
    date_col="date",
    data_options=None
)

Joins lagged values of selected variables from one data frame (‘new_data’) into another (‘original_data’), based on date ranges defined by ‘min_lag’ and ‘max_lag’. Unlike ‘add_lagged_columns’, this function supports joining across data frames with different date grids (e.g. monthly source data into quarterly target data). All columns in ‘new_data’ besides ‘id_keys’ and the date column are lagged and joined under their original names.

Parameters

original_data: pl.DataFrame

Target panel data. The date column must be of dtype ‘pl.Date’ or ‘pl.Datetime’.

new_data: pl.DataFrame

Source variables to lag and merge. All columns besides ‘id_keys’ and the date column will be lagged and joined.

id_keys: list of str or str

Identifier column(s) shared by both frames.

min_lag: int, str, datetime.timedelta, or pd.DateOffset

Lower lag bound (inclusive).

max_lag: int, str, datetime.timedelta, or pd.DateOffset

Upper lag bound (inclusive).

ff_adjustment: bool = False

If True, keeps only the last observation per identifier and year in ‘new_data’ before lagging (Fama-French convention). Defaults to False.

date_col: str = "date"

Name of the date column. Defaults to ‘date’.

data_options: dict = None
Column-name mapping (see ‘data_options’). The ‘date’ element is used to identify the date column. Uses the ‘data_options’ default when None: ‘date’ -> ‘date’.

Returns

pl.DataFrame
‘original_data’ with all columns from ‘new_data’ appended as lagged values (keeping their original names).

Examples

import numpy as np
import polars as pl
import datetime as dt
from tidyfinance import join_lagged_values
rng = np.random.default_rng(42)
dates = pl.date_range(
    dt.date(2020, 1, 1), dt.date(2020, 6, 1), "1mo", eager=True
)
df1 = pl.DataFrame({
    'id': [1] * 6 + [2] * 6,
    'date': dates.to_list() * 2,
})
df2 = df1.with_columns(
    x=pl.Series(rng.standard_normal(len(df1)))
)
join_lagged_values(
    original_data=df1,
    new_data=df2,
    id_keys='id',
    min_lag='1mo',
    max_lag='3mo',
)