Skip to main content

๐ŸชŸ Window Functions

๐Ÿ“š What You'll Learn

By the end of this lesson, you will be able to:

  • Explain what a "window" is and how window functions differ from a plain aggregation
  • Use .rolling() for fixed-size sliding windows (moving averages, rolling std, and more)
  • Use .expanding() for cumulative, growing-from-the-start windows
  • Use .ewm() for exponentially weighted windows that favor recent observations
  • Apply any aggregation โ€” mean, sum, min/max, std, custom โ€” inside a window
  • Combine windows with .rank(), .shift(), and cumulative functions for real feature engineering

โฑ๏ธ Estimated Time: 45โ€“60 minutes

๐ŸŽฏ Project: Build a moving-metrics toolkit โ€” rolling averages and volatility, expanding running totals, EWM smoothing, and window-based rankings โ€” on a numeric series.

Compute statistics over a moving slice of your data instead of the whole column at once.

๐ŸŒŸ Statistics That Slide

A normal aggregation collapses an entire column into a single number: one mean, one sum, one max. A window function is different โ€” it computes that statistic over a moving slice of the data and returns a value for every position. The result lines up with your original rows, so you can attach a 7-day moving average, a running total, or a smoothed trend right next to the raw numbers.

Windows come in three flavors: a rolling window is a fixed-size frame that slides along; an expanding window starts small and grows to include everything seen so far; and an exponentially weighted window keeps all the history but gives recent points more say. Master these three and you can describe how any measurement behaves over its sequence, not just on average.

๐Ÿ”Ž The Three Window Types

graph TD A[A numeric Series] --> B[.rolling(n)
fixed sliding frame] A --> C[.expanding()
grows from the start] A --> D[.ewm()
weighted toward recent] B --> E[same-length result
aligned to each row] C --> E D --> E style A fill:#f9f,stroke:#333,stroke-width:2px style E fill:#9f9,stroke:#333,stroke-width:2px

๐ŸชŸ .rolling()

A fixed-size window slides across the series, one step at a time.

# 3-period moving average
s.rolling(window=3).mean()

# Rolling std (volatility)
s.rolling(window=5).std()

# Require full window before output
s.rolling(3, min_periods=3).sum()

๐Ÿ“ˆ .expanding()

The window starts at the first row and grows to include every prior value.

# Running (cumulative) mean
s.expanding().mean()

# Running maximum so far
s.expanding().max()

# Same idea as cumsum for sums
s.expanding().sum()

โš–๏ธ .ewm()

Every past point counts, but weights decay geometrically toward the present.

# Exponentially weighted mean
s.ewm(alpha=0.3).mean()

# Specify by span (like a 10-period EMA)
s.ewm(span=10).mean()

# Weighted volatility
s.ewm(span=10).std()

๐ŸŽฎ Interactive Sliding Window

Slide the frame across the data and choose an aggregation to see exactly what a rolling window computes at each step.

3 step 1
window result: -

๐Ÿ“‰ Rolling Average in Action

The classic use of a rolling window is smoothing. The faint line is raw, noisy data; the bold line is its 10-period rolling mean โ€” the same signal with the jitter removed.

# Smooth noisy data with a moving average
df['smooth'] = df['value'].rolling(window=10).mean()

# Bollinger-style bands from a rolling window
roll = df['value'].rolling(20)
df['upper'] = roll.mean() + 2 * roll.std()
df['lower'] = roll.mean() - 2 * roll.std()

๐Ÿ“ˆ Expanding Windows

An expanding window answers "what is the statistic using everything up to now?" Each bar below is a data point; the number above it is the running mean of all points seen so far โ€” notice how it stabilizes as more data accumulates.

# Running average that never forgets
df['running_mean'] = df['value'].expanding().mean()

# Running total (equivalent to cumsum)
df['running_total'] = df['value'].expanding().sum()

# Best value seen so far
df['record'] = df['value'].expanding().max()

โš–๏ธ Exponentially Weighted Windows

An EWM keeps the whole history but weights it with a decaying factor. The bars show how much influence each past observation (t, t-1, t-2, โ€ฆ) has on today's value โ€” recent points dominate, distant ones fade.

# alpha controls how fast old data fades
df['ewm'] = df['value'].ewm(alpha=0.3).mean()

# Equivalent parameterizations
df['value'].ewm(span=10).mean()       # like a 10-period EMA
df['value'].ewm(halflife=5).mean()    # weight halves every 5 steps

๐Ÿ… Ranking Within Windows

Rankings are close cousins of window functions โ€” they compare each value against the others in a group. Below, the same six scores are ranked several ways at once.

NameScoreRank Dense RankPercentileQuartile
# Standard competition rank (ties share a rank, gaps follow)
df['rank'] = df['score'].rank(ascending=False)

# Dense rank (no gaps after ties)
df['dense'] = df['score'].rank(method='dense', ascending=False)

# Percentile rank in [0, 1]
df['pct'] = df['score'].rank(pct=True)

๐Ÿ’ก Pro Tips for Window Functions

โš ๏ธ Common Pitfalls to Avoid

๐Ÿ“‹ Quick Reference

Window Types:

Common Window Functions:

Shift Operations:

Ranking Functions:

Cumulative Functions:

๐Ÿ““ Learning Journal

Keep a learning journal โ€” digital or physical. After this lesson, take a few minutes to write down:

  • Key concepts you learned
  • Techniques that clicked for you
  • Questions or confusion points to revisit
  • Ideas you want to try
  • Your progress and feelings about learning this

โœ๏ธ This lesson's prompt: Rolling, expanding, and EWM each "remember" the past differently. Pick a measurement you track over time and decide which window fits: would a fixed recent window, a running-since-the-start view, or a recency-weighted average tell the most honest story โ€” and why?

๐Ÿ“ Lesson Summary

๐ŸŽ“ Key Takeaways

  • Window functions compute a statistic over a moving slice and return a value aligned to every row โ€” unlike an aggregation that returns one number.
  • .rolling(n) uses a fixed sliding frame, .expanding() grows from the start, and .ewm() weights recent points more heavily.
  • Any aggregation โ€” mean, sum, std, min/max, or a custom .apply() โ€” can run inside a window.
  • Mind the warm-up NaNs, sort first, and avoid look-ahead bias when building features for prediction.

๐ŸŽ‰ What You've Accomplished

You can now describe how a measurement behaves across its own sequence โ€” smoothing noise with rolling means, tracking running totals with expanding windows, emphasizing recency with EWM, and ranking values within a window. These are the building blocks of time-aware feature engineering.

โ“ Common Questions at This Stage

When should I use expanding() instead of rolling()?

Use rolling() when only the recent past matters and older data should drop out of the window (e.g. a 30-day moving average). Use expanding() when every observation so far should keep contributing โ€” running totals, cumulative averages, or "best/worst to date" metrics.

What does alpha mean in .ewm()?

alpha (between 0 and 1) is the smoothing factor: it is the weight given to the most recent point, with every older point weighted by alphaยท(1-alpha)^k. A larger alpha reacts faster to change; a smaller alpha smooths more. You can also specify the decay indirectly via span or halflife.

Why are the first few values of my rolling result NaN?

A rolling window needs a full window worth of observations before it can produce a value, so the first window-1 results are NaN. Set min_periods to a smaller number if you are willing to compute on a partial window at the start.

๐Ÿ”ญ Looking Ahead

Window operations can get expensive on large data. Next you'll focus on performance โ€” vectorization, efficient dtypes, and avoiding slow .apply() calls โ€” so your rolling and grouped computations stay fast at scale.

โœ… Before the Next Lesson

๐ŸŒŸ Encouragement for the Journey

Window functions are the quiet workhorses of real analytics โ€” the moving averages behind every dashboard and the smoothed lines behind every trend. Once you can slide a window in your head, a whole layer of data behavior becomes visible. Keep sliding!