lagging.add_lagged_columns()
Append lagged columns to a data frame via a join-based approach.
Usage
lagging.add_lagged_columns(
data,
cols,
lag,
max_lag=None,
by=None,
drop_na=False,
ff_adjustment=False,
date_col="date",
data_options=None
)When ‘lag == max_lag’ (the default), an equi-join is used: source dates are shifted forward by ‘lag’ and matched exactly. When ‘lag < max_lag’, an inequality join is used: for each row, the most recent source value within the window ‘[date - max_lag, date - lag]’ is selected.
The combination of ‘by’ and the date column must be unique in ‘data’. If ‘by’ is None, dates alone must be unique.
Parameters
data: pl.DataFrame-
DataFrame containing the variables to lag. The date column must be of dtype ‘pl.Date’ or ‘pl.Datetime’.
cols: list of str or str-
Names of the columns to lag. Each column produces a new column suffixed with ’_lag’.
lag: int, str, datetime.timedelta, or pd.DateOffset-
Minimum lag (inclusive) to apply, e.g. ‘“1mo”’. An int is interpreted as days; strings are polars offset strings.
max_lag: int, str, datetime.timedelta, or pd.DateOffset = None-
Maximum lag (inclusive) to apply. Defaults to ‘lag’ (exact lag).
by: list of str or str = None-
Grouping column(s) (e.g. a stock identifier). Lagged values are matched within groups. Defaults to None.
drop_na: bool = False-
If True, missing values in the source columns are excluded before matching, so the lookup skips over missing observations. Applied independently per column. Defaults to False.
ff_adjustment: bool = False-
If True, only the last observation per year (within each group defined by ‘by’) is retained as a source for lagged values, following Fama-French conventions for annual accounting data. Defaults to False.
date_col: str = "date"-
Name of the date column. Defaults to ‘date’.
data_options: dict = None- Column-name mapping (see ‘data_options’). The ‘date’ element is used to specify the date column. Uses the ‘data_options’ default when None: ‘date’ -> ‘date’.
Returns
pl.DataFrame- Data frame with the same rows as ‘data’ and new columns appended, each suffixed with ’_lag’. Unmatched rows receive null in the lagged columns.
Examples
import numpy as np
import polars as pl
import datetime as dt
from tidyfinance import add_lagged_columns
rng = np.random.default_rng(42)
dates = pl.date_range(
dt.date(2023, 1, 1), dt.date(2023, 10, 1), "1mo", eager=True
)
data = pl.DataFrame({
'permno': [1] * 10 + [2] * 10,
'date': dates.to_list() * 2,
'size': rng.uniform(100, 200, 20),
'bm': rng.uniform(0.5, 1.5, 20),
})
# Exact lag: each row gets the value from exactly 2 months earlier
add_lagged_columns(data, cols=['size', 'bm'], lag='2mo', by='permno')
# Window lag: most recent value from 2 to 4 months earlier
add_lagged_columns(
data, cols='size', lag='2mo', max_lag='4mo', by='permno'
)