lagging.compute_rolling_value()

Compute a rolling value by period.

Usage

Source

lagging.compute_rolling_value(
    data, f, period="month", periods=12, min_obs=None, data_options=None
)

Applies an arbitrary summary function over rolling time-period windows. Each window spans ‘periods’ units of ‘period’ (e.g., 12 months). Before calling ‘f’, rows with any missing values are dropped from the window; if fewer than ‘min_obs’ rows remain, the result is NaN instead.

Parameters

data: pl.DataFrame

Data frame with a date column named according to ‘data_options[date]’ (default ‘date’). The column must be of dtype ‘pl.Date’ or ‘pl.Datetime’.

f: callable

Function applied to each window. Receives the window slice (complete cases only) as a data frame of the active backend type (‘pd.DataFrame’ under the default ‘pandas’ backend, ‘pl.DataFrame’ under the ‘polars’ backend) and must return a single scalar value.

period: str = "month"

Calendar period unit for the rolling windows. One of ‘month’, ‘quarter’, or ‘year’.

periods: int = 12

Number of periods to include in the rolling window.

min_obs: int = None

Minimum number of non-missing rows required per window. Defaults to ‘periods’.

data_options: dict = None
Column-name mapping (see ‘data_options’). The ‘date’ element is used to specify the date column. Uses the ‘data_options’ default when None: ‘date’ -> ‘date’.

Returns

np.ndarray
Numeric vector aligned with the rows of ‘data’.

Examples

import numpy as np
import polars as pl
import datetime as dt
from tidyfinance import compute_rolling_value
rng = np.random.default_rng(42)
df = pl.DataFrame({
    'date': pl.date_range(
        dt.date(2020, 1, 1), dt.date(2021, 12, 1), '1mo', eager=True
    ),
    'value': rng.standard_normal(24),
})
df = df.with_columns(
    rolling_sd=pl.Series(compute_rolling_value(
        df,
        f=lambda x: x['value'].std(ddof=1),
        period='month',
        periods=4,
        min_obs=2,
    ))
)