Skip to content

Concepts

Understanding how LOWESS works and when to use it.

LOWESS (Locally Weighted Scatterplot Smoothing) is a nonparametric regression method that fits smooth curves through scatter plots without assuming a global functional form.

Unlike parametric methods (linear regression, polynomial fitting), LOWESS adapts locally to the data structure, making it ideal for:

  • Exploratory data analysis — Discover patterns without assumptions
  • Trend estimation — Extract signals from noisy time series
  • Baseline correction — Remove systematic effects in spectroscopy
  • Genomic smoothing — Smooth methylation, ChIP-seq, or expression data

LOWESS Smoothing Concept

LOWESS fits local weighted regressions at each point, using a focused local window around each evaluation point

For each point in your data, LOWESS:

  1. Selects neighbors — Choose the nearest points (controlled by fraction)
  2. Assigns weights — Closer points get higher weights (using a kernel function)
  3. Fits locally — Perform weighted least squares regression
  4. Extracts value — Use the fitted value as the smoothed estimate
  5. Iterates (optional) — Reweight points based on residuals to reduce outlier influence

The fraction (also called bandwidth or span) is the most important parameter. It controls what proportion of data is used for each local fit.

Fraction Effect

Small fraction vs large fraction — bandwidth controls how closely the fit follows local structure

FractionEffectWhen to Use
0.1–0.3Fine detail, follows data closelyRapidly changing signals
0.3–0.5Balanced smoothingMost applications
0.5–0.7Heavy smoothingNoisy data, trend extraction
0.7–1.0Very smoothStrong noise, global trends

Standard LOWESS is sensitive to outliers. Robustness iterations downweight points with large residuals:

Robustness Effect

Non-robust LOWESS (iterations=0) vs robust LOWESS — outlier influence is suppressed through iterative reweighting

IterationsEffectWhen to Use
0No robustness (fastest)Clean data, speed-critical
1–3Moderate robustnessMost applications
4–6Strong robustnessData with outliers
7+Very strongHeavy contamination

Intervals

Confidence intervals (narrow, mean curve uncertainty) vs Prediction intervals (wide, new-point uncertainty)

Interval TypeWhat It RepresentsWidth
ConfidenceUncertainty in the mean curveNarrow
PredictionUncertainty for new observationsWide
  • Use confidence intervals to show where the true trend likely lies
  • Use prediction intervals to show where new data points might fall

Choose the right mode based on your use case:

ModeUse CaseMemoryFeatures
BatchComplete datasetsEntire datasetAll features
StreamingLarge files (>100K points)One chunkResiduals, robustness
OnlineReal-time dataFixed windowIncremental updates

SituationMode
Data fits in memory; needs intervals or CVBatch
Data too large for memory or arrives in chunksStreaming
Data arrives point-by-point in real timeOnline

FeatureLOWESSPolynomial RegressionMoving Average
No parametric assumptions✓✗✓
Adapts to local structure✓✗Partial
Robust to outliers✓✗✗
Uncertainty estimates✓✓✗
Handles irregular sampling✓✓✗