lagging.add_lagged_columns()

Append lagged columns to a data frame via a join-based approach.

Usage

Source

lagging.add_lagged_columns(
    data,
    cols,
    lag,
    max_lag=None,
    by=None,
    drop_na=False,
    ff_adjustment=False,
    date_col="date",
    data_options=None
)

When ‘lag == max_lag’ (the default), an equi-join is used: source dates are shifted forward by ‘lag’ and matched exactly. When ‘lag < max_lag’, an inequality join is used: for each row, the most recent source value within the window ‘[date - max_lag, date - lag]’ is selected.

The combination of ‘by’ and the date column must be unique in ‘data’. If ‘by’ is None, dates alone must be unique.

Parameters

data: pl.DataFrame

DataFrame containing the variables to lag. The date column must be of dtype ‘pl.Date’ or ‘pl.Datetime’.

cols: list of str or str

Names of the columns to lag. Each column produces a new column suffixed with ’_lag’.

lag: int, str, datetime.timedelta, or pd.DateOffset

Minimum lag (inclusive) to apply, e.g. ‘“1mo”’. An int is interpreted as days; strings are polars offset strings.

max_lag: int, str, datetime.timedelta, or pd.DateOffset = None

Maximum lag (inclusive) to apply. Defaults to ‘lag’ (exact lag).

by: list of str or str = None

Grouping column(s) (e.g. a stock identifier). Lagged values are matched within groups. Defaults to None.

drop_na: bool = False

If True, missing values in the source columns are excluded before matching, so the lookup skips over missing observations. Applied independently per column. Defaults to False.

ff_adjustment: bool = False

If True, only the last observation per year (within each group defined by ‘by’) is retained as a source for lagged values, following Fama-French conventions for annual accounting data. Defaults to False.

date_col: str = "date"

Name of the date column. Defaults to ‘date’.

data_options: dict = None
Column-name mapping (see ‘data_options’). The ‘date’ element is used to specify the date column. Uses the ‘data_options’ default when None: ‘date’ -> ‘date’.

Returns

pl.DataFrame
Data frame with the same rows as ‘data’ and new columns appended, each suffixed with ’_lag’. Unmatched rows receive null in the lagged columns.

Examples

import numpy as np
import polars as pl
import datetime as dt
from tidyfinance import add_lagged_columns
rng = np.random.default_rng(42)
dates = pl.date_range(
    dt.date(2023, 1, 1), dt.date(2023, 10, 1), "1mo", eager=True
)
data = pl.DataFrame({
    'permno': [1] * 10 + [2] * 10,
    'date': dates.to_list() * 2,
    'size': rng.uniform(100, 200, 20),
    'bm': rng.uniform(0.5, 1.5, 20),
})
# Exact lag: each row gets the value from exactly 2 months earlier
add_lagged_columns(data, cols=['size', 'bm'], lag='2mo', by='permno')
# Window lag: most recent value from 2 to 4 months earlier
add_lagged_columns(
    data, cols='size', lag='2mo', max_lag='4mo', by='permno'
)