Skip to content

StreamingLowess API

See also: fastLowess

  • Dataset >100,000 points
  • Memory-constrained environments
  • Batch processing pipelines

The StreamingLowess class processes data in chunks, suitable for very large datasets or streaming applications.

Constructor:

const { StreamingLowess } = require('fastlowess');
const stream = new StreamingLowess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });
const x = Float64Array.from({ length: 10 }, (_, i) => i);
const y = Float64Array.from({ length: 10 }, (_, i) => i * 0.5);
stream.process_chunk(x, y);
const result = stream.finalize();
console.log("Smoothed", result.y.length, "points via streaming");
Smoothed 10 points via streaming
  • options: An object containing StreamingSmoothOptions fields (a subset of the Batch LowessOptions fields — see below).
  • streamingOptions: An object containing StreamingOptions fields.

Feeds one chunk of data into the model. Each chunk is fit together with the trailing overlap points buffered from the previous call, then only the points that are fully resolved are returned — the tail of the chunk (the next overlap points) is held back internally, since it will be refit once the following chunk arrives and its estimate reconciled via merge_strategy. This is what lets the adapter process a dataset far larger than memory allows, one bounded-size chunk at a time, without ever materializing the whole dataset at once.

const { StreamingLowess } = require('fastlowess');
const n = 100;
const x = Float64Array.from({ length: n }, (_, i) => i * 2 * Math.PI / (n - 1));
const y = Float64Array.from(x, xi => Math.sin(xi) + 0.1);
const stream = new StreamingLowess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });
const partialResult = stream.process_chunk(x.slice(0, 50), y.slice(0, 50));
console.log("Fraction used:", partialResult.fraction_used);
Fraction used: 0.5

Flushes the overlap points still buffered from the last process_chunk() call. Because each call withholds its tail until the next chunk arrives to resolve it, the final chunk’s tail would never be emitted otherwise — always call finalize() once after the last chunk to retrieve it.

const { StreamingLowess } = require('fastlowess');
const n = 100;
const x = Float64Array.from({ length: n }, (_, i) => i * 2 * Math.PI / (n - 1));
const y = Float64Array.from(x, xi => Math.sin(xi) + 0.1);
const stream = new StreamingLowess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });
stream.process_chunk(x.slice(0, 50), y.slice(0, 50));
stream.process_chunk(x.slice(50), y.slice(50));
const finalResult = stream.finalize();
console.log("Fraction used:", finalResult.fraction_used);
Fraction used: 0.5
FieldTypeDefaultDescription
fractionnumber0.67Smoothing fraction (bandwidth)
iterationsnumber3Number of robustifying iterations
weight_functionstring"tricube"Weight function name
robustness_methodstring"bisquare"Robustness method name
deltanumberNaNInterpolation distance (NaN auto-sets it to 0.0 in Streaming, i.e. interpolation disabled)
zero_weight_fallbackstring"use_local_mean"Zero-weight handling
boundary_policystring"extend"Boundary handling policy
scaling_methodstring"mad"Residual scaling method
auto_convergenumbernullAuto-convergence tolerance
missingstring"error"Policy for non-finite (NaN/Inf) values in each chunk
chunk_sizenumber5000Data chunk size
overlapnumberchunk_size / 10Overlap between chunks
merge_strategystring"weighted_average"Strategy for blending overlap regions
parallelbooleantrueEnable parallel execution
outputsstring[][]Select se, diagnostics, residuals, weights, and/or derivative
intervalsobjectnullGrouped confidence, prediction, and per-chunk bootstrap options
seednumbernullReproducible bootstrap draws for each combined chunk; must be an integer from 0 through Number.MAX_SAFE_INTEGER

Cross-validation, GPU backend, custom_weights, and the "sorted" output are Batch-only and not available here; see fastLowess for those. Standard errors and confidence/prediction intervals are computed per combined chunk (including the previous overlap), then blended across overlap regions via merge_strategy like y/derivative are; they are local chunk intervals, not whole-stream intervals.

fraction is the most important parameter: it controls the size of the local neighbourhood used at each point.

RangeEffectUse case
0.1-0.3Fine detailRapidly changing signals
0.3-0.5BalancedGeneral purpose
0.5-0.7Heavy smoothingNoisy data
0.7-1.0Very smoothTrend extraction

iterations controls robustness to outliers, at the cost of speed.

ValueEffectPerformance
0No robustnessFastest
1-3ModerateRecommended
4-6StrongContaminated data
7+Very strongHeavy outliers

See: Weight Functions

  • "tricube" (default)
  • "epanechnikov"
  • "gaussian"
  • "uniform" (alias: "boxcar")
  • "biweight" (alias: "bisquare")
  • "triangle" (alias: "triangular")
  • "cosine"

See: Robustness

  • "bisquare" (default; alias: "biweight")
  • "huber"
  • "talwar"

Points within delta of each other on the x-axis share the same local fit instead of each computing its own regression — an interpolation shortcut that trades a small amount of accuracy for a large speedup on dense, evenly-spaced data. NaN (default) auto-sets it to 0 in Streaming mode, i.e. interpolation is disabled and every point is fit exactly.

Behavior when all neighborhood weights are zero:

OptionBehavior
"use_local_mean" (default; aliases: "local_mean", "mean")Use the mean of the neighborhood
"return_original" (alias: "original")Return the original y value
"return_none" (alias: "none")Return NaN

See: Boundary Handling

  • "extend" (default; alias: "pad")
  • "reflect" (alias: "mirror")
  • "zero"
  • "noboundary" (alias: "none")

See: Scaling Methods

  • "mad" (default; alias: "median_absolute_deviation")
  • "mar" (alias: "median_absolute_residual")
  • "mean" (alias: "mean_absolute_residual")

See: Robustness

Convergence tolerance for early stopping of robustness iterations. null (default) disables early stopping.

Policy for handling non-finite (NaN/Inf) values within each chunk:

OptionBehavior
"error" (default)Throw an error if any value in the chunk is non-finite
"drop"Silently remove rows where x or y is non-finite before merging the chunk with the overlap buffer

Note: A length mismatch between x and y always errors, even under "drop".

Number of points processed per chunk. Larger chunks reduce per-chunk overhead and give each local fit more surrounding context, at the cost of higher peak memory; smaller chunks bound memory tightly but increase the fraction of points that fall in overlap regions. A good starting point is balancing available memory against how much processing overhead per chunk is acceptable — match it to your file-read buffer or message-batch size to avoid unnecessary copying.

Number of points retained from the previous chunk as context, so the neighbourhood at chunk boundaries isn’t artificially truncated. Points inside the overlap zone are fitted twice (once by each chunk) and reconciled via merge_strategy. A good starting point is 10–20% of chunk_size: too little overlap causes visible boundary artefacts, while too much wastes computation refitting the same points twice.

  • null (default) — computes chunk_size / 10, clamped to at least 1 and less than chunk_size
  • Any integer >= 1 and < chunk_size

See: Merge Strategies

StrategyAliasBehavior
"weighted_average" (default)"weighted"Distance-weighted blend
"average""mean"Average overlapping values
"take_first""first"Keep left chunk values
"take_last""last"Keep right chunk values

Enable multi-threaded execution via Rayon.

  • true (default) — parallelizes the local regression fits across CPU cores
  • false — forces single-threaded execution

See: Intervals

Computes standard errors per chunk the same way Batch does, then merges the overlap region across chunk boundaries the same way y/derivative are, via merge_strategy.

See: Diagnostics

Select "diagnostics" to include a Diagnostics object (RMSE, MAE, R², residual_sd) in the result. effective_df/aic/aicc require per-chunk hat-matrix leverage to be threaded into the cumulative diagnostics computation across chunk boundaries, which isn’t currently done (even with "se" or intervals selected), so they’re always null here.

Include per-point residuals (y - fitted) in the result.

Include the final per-point robustness weights (from the last robustness iteration) in the result.

Each point’s local WLS fit already computes a slope internally; this exposes that per-point slope (rate of change of the smoothed curve) in LowessResult.derivative at effectively no extra computation cost. Derivative values in the overlap region are merged across chunk boundaries the same way y is, via merge_strategy.

See: Intervals

An object such as { confidence: 0.90, prediction: 0.99, bootstrap: 200 }, populating result.confidence_lower/result.confidence_upper and result.prediction_lower/result.prediction_upper. The coverage levels are independent. Computed per combined chunk and merged across overlap boundaries via merge_strategy. bootstrap (at least 2) refits each combined chunk from resampled residuals; with parallel: true, refits run concurrently.

Seeds bootstrap draws. Each combined chunk restarts from the same seed. It does not enable bootstrap by itself. Seeds must be finite, non-negative safe integers no greater than Number.MAX_SAFE_INTEGER; fractional, negative, and unsafe values throw. 0 is valid.

Returned by process_chunk() and finalize().

FieldTypeDescription
xFloat64Arrayx values (same order as input)
yFloat64ArraySmoothed y values
fraction_usednumberFraction used
iterations_usednumber | nullRobustness iterations actually performed
standard_errorsFloat64Array | nullPer-point standard errors (if "se" or any interval was set)
confidence_lowerFloat64Array | nullLower confidence bounds (if intervals.confidence was set)
confidence_upperFloat64Array | nullUpper confidence bounds (if intervals.confidence was set)
prediction_lowerFloat64Array | nullLower prediction bounds (if intervals.prediction was set)
prediction_upperFloat64Array | nullUpper prediction bounds (if intervals.prediction was set)
residualsFloat64Array | nullResiduals (if "residuals" was requested)
robustness_weightsFloat64Array | nullRobustness weights (if "weights" was requested)
cv_scoresFloat64Array | nullAlways null (Batch only)
diagnosticsDiagnostics | nullFit metrics (if "diagnostics" was requested)
derivativeFloat64Array | nullPer-point local fit derivative/slope (if "derivative" was requested)
FieldTypeDescription
rmsenumberRoot Mean Squared Error
maenumberMean Absolute Error
r_squarednumberR-squared
residual_sdnumberCumulative sample SD of emitted residuals
effective_dfnumber | nullAlways null (cumulative diagnostics don’t integrate per-chunk leverage; Batch only)
aicnumber | nullAlways null (requires effective_df; Batch only)
aiccnumber | nullAlways null (requires effective_df; Batch only)