Skip to content

StreamingLoess API

See also: fastLoess

  • Dataset >100,000 points
  • Memory-constrained environments
  • Batch processing pipelines

The StreamingLoess class processes data in chunks, suitable for very large datasets or streaming applications.

Constructor:

const { StreamingLoess } = require('fastloess-wasm');
const stream = new StreamingLoess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });
console.log("typeof process_chunk:", typeof stream.process_chunk);
typeof process_chunk: function
  • options: An object containing StreamingSmoothOptions fields (a subset of the Batch LoessOptions fields — see below).
  • streamingOptions: An object containing StreamingOptions fields.

process_chunk_weighted(x, y, weights) accepts one finite non-negative case weight per observation and preserves weights across buffered overlap points.

Processes a chunk of data. Returns partial results.

const { StreamingLoess } = require('fastloess-wasm');
const n = 100;
const x = Float64Array.from({ length: n }, (_, i) => i * 2 * Math.PI / (n - 1));
const y = Float64Array.from(x, xi => Math.sin(xi) + 0.1);
const stream = new StreamingLoess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });
const partialResult = stream.process_chunk(x.slice(0, 50), y.slice(0, 50));
console.log("Fraction used:", partialResult.fraction_used);
Fraction used: 0.5

Finalizes the smoothing process and returns any remaining buffered results.

const { StreamingLoess } = require('fastloess-wasm');
const n = 100;
const x = Float64Array.from({ length: n }, (_, i) => i * 2 * Math.PI / (n - 1));
const y = Float64Array.from(x, xi => Math.sin(xi) + 0.1);
const stream = new StreamingLoess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });
stream.process_chunk(x.slice(0, 50), y.slice(0, 50));
stream.process_chunk(x.slice(50), y.slice(50));
const finalResult = stream.finalize();
console.log("Fraction used:", finalResult.fraction_used);
Fraction used: 0.5
FieldTypeDefaultDescription
fractionnumber0.67Smoothing fraction (bandwidth)
iterationsnumber3Number of robustifying iterations
weight_functionstring"tricube"Weight function name
robustness_methodstring"bisquare"Robustness method name
degreestring"linear"Polynomial degree of local fit
dimensionsnumber1Number of predictor dimensions
distance_metricstring"normalized"Distance metric; use "minkowski:p" for custom p
weighted_metric_weightsnumber[]nullPer-dimension weights (used when distance_metric = "weighted")
surface_modestring"interpolation"Surface computation mode
cellnumbernullCell size for interpolation grid (smaller → more vertices, higher accuracy)
interpolation_verticesnumbernullNumber of interpolation vertices
zero_weight_fallbackstring"use_local_mean"Zero-weight handling strategy
boundary_policystring"extend"Boundary handling policy
boundary_degree_fallbackbooleannullFall back to lower polynomial degree at boundaries when higher degrees fail
scaling_methodstring"mad"Residual scaling method
auto_convergenumbernullAuto-convergence tolerance
missingstring"error"Policy for non-finite (NaN/Inf) values in each chunk
parallelbooleantrueEnable parallel execution
outputsstring[][]Optional fields: "diagnostics", "residuals", "weights", "gradient" (or "derivative"), "se"
intervals{ confidence?: number; prediction?: number }disabledGrouped confidence and prediction coverage levels.

"sorted" output and cross-validation are Batch-only and not available here; see fastLoess.

FieldTypeDefaultDescription
chunk_sizenumber5000Data chunk size
overlapnumberchunk_size / 10Overlap between chunks
merge_strategystring"weighted_average"Strategy for blending overlap regions

fraction is the most important parameter: it controls the size of the local neighbourhood used at each point.

RangeEffectUse case
0.1-0.3Fine detailRapidly changing signals
0.3-0.5BalancedGeneral purpose
0.5-0.7Heavy smoothingNoisy data
0.7-1.0Very smoothTrend extraction

iterations controls robustness to outliers, at the cost of speed.

ValueEffectPerformance
0No robustnessFastest
1-3ModerateRecommended
4-6StrongContaminated data
7+Very strongHeavy outliers

See: Weight Functions

  • "tricube" (default)
  • "epanechnikov"
  • "gaussian"
  • "uniform" (alias: "boxcar")
  • "biweight" (alias: "bisquare")
  • "triangle" (alias: "triangular")
  • "cosine"

See: Robustness

  • "bisquare" (default; alias: "biweight")
  • "huber"
  • "talwar"

See: Polynomial Degree

  • "constant" or "0" (degree 0)
  • "linear" or "1" (default, degree 1)
  • "quadratic" or "2" (degree 2)
  • "cubic" or "3" (degree 3)
  • "quartic" or "4" (degree 4)

See: Multivariate LOESS

Number of predictor dimensions. Set to match the number of columns in a multivariate x array.

  • Any integer >= 1; 1 (default) is univariate

See: Multivariate LOESS

  • "normalized" (default — scales each dimension by its range; alias: "norm")
  • "euclidean" (alias: "euclid")
  • "manhattan" (alias: "l1")
  • "chebyshev" (alias: "linf")
  • "minkowski" (use "minkowski:p" string for custom exponent, e.g. "minkowski:3")
  • "weighted" plus weighted_metric_weights for per-dimension scaling (alias: "weighted_euclidean")

See: Multivariate LOESS

Per-dimension weights, one per dimension declared in dimensions. Only used when distance_metric = "weighted"; setting distance_metric = "weighted" without providing this raises an error.

  • null (default) — has no effect unless distance_metric = "weighted" is set
  • A number[] of per-dimension weights, required when distance_metric = "weighted"

See: Polynomial Degree

Controls whether the local polynomial is evaluated at every query point or at a sparser grid of anchor vertices with Hermite cubic interpolation in between.

ModeBehaviorSpeedAccuracy
"interpolation" (default)Evaluate at vertices, interpolate betweenFasterSlight approximation
"direct"Evaluate at every query pointSlowerFull precision

Cell size for the interpolation grid, as a fraction of the data range. Smaller values place more vertices (denser grid), improving accuracy at the cost of speed. Only applies when surface_mode = "interpolation".

  • null (default) — uses the library default (0.2)
  • Any number in (0, 1]

Caps the maximum number of interpolation vertices, overriding the count implied by cell. Only applies when surface_mode = "interpolation".

  • null (default) — uses the library default (no explicit cap)
  • Any integer >= 1

Behavior when all neighborhood weights are zero:

OptionBehavior
"use_local_mean" (default; aliases: "local_mean", "mean")Use the mean of the neighborhood
"return_original" (alias: "original")Return the original y value
"return_none" (alias: "none")Return NaN

See: Boundary Handling

  • "extend" (default; alias: "pad")
  • "reflect" (alias: "mirror")
  • "zero"
  • "noboundary" (alias: "none")

Whether to reduce the polynomial degree at boundary vertices when the requested degree can’t be fit there (e.g., not enough neighbours). Only applies when surface_mode = "interpolation".

  • null (default) — uses the library default (enabled)
  • true — falls back to a lower degree at boundaries
  • false — raises an error instead of silently falling back

See: Scaling Methods

  • "mad" (default; alias: "median_absolute_deviation")
  • "mar" (alias: "median_absolute_residual")
  • "mean" (alias: "mean_absolute_residual")

See: Robustness

Convergence tolerance for early stopping of robustness iterations. null (default) disables early stopping.

Policy for handling non-finite (NaN/Inf) values within each chunk:

OptionBehavior
"error" (default)Throw an error if any value in the chunk is non-finite
"drop"Silently remove rows where any x dimension or y is non-finite before merging the chunk with the overlap buffer

Note: A length mismatch between x and y always throws, even under "drop".

Number of points processed per call to process_chunk(). Larger chunks reduce per-chunk overhead and give each local fit more surrounding context, at the cost of higher peak memory; smaller chunks bound memory tightly but increase the fraction of points that fall in overlap regions. A good starting point is balancing available memory against how much processing overhead per chunk is acceptable — match it to your file-read buffer or message-batch size to avoid unnecessary copying.

Number of points retained from the previous chunk as context, so the neighbourhood at chunk boundaries isn’t artificially truncated. Points inside the overlap zone are fitted twice (once by each chunk) and reconciled via merge_strategy. A good starting point is 10–20% of chunk_size: too little overlap causes visible boundary artefacts, while too much wastes computation refitting the same points twice.

  • null (default) — computes chunk_size / 10, clamped to at least 1 and at most chunk_size - 10
  • Any integer >= 1 and < chunk_size

See: Merge Strategies

StrategyAliasBehavior
"weighted_average" (default)"weighted"Distance-weighted blend
"average""mean"Average overlapping values
"take_first""first"Keep left chunk values
"take_last""last"Keep right chunk values

Merge Strategies

Enable multi-threaded execution via the Rayon-based web worker pool.

  • true (default) — parallelizes the local regression fits
  • false — forces single-threaded execution (useful for benchmarking or deterministic profiling)

Include standard errors in the result (LoessResult.standard_errors), computed per chunk and merged across overlap boundaries via merge_strategy.

Include a Diagnostics object (RMSE, MAE, R², residual_sd) in the result. Streaming residual_sd is the cumulative sample standard deviation of emitted residuals. effective_df/aic/aicc require standard errors, which are Batch-only, so they’re always null here.

Include per-point residuals (y - fitted) in the result.

Include the final per-point robustness weights (from the last robustness iteration) in the result.

Each local polynomial fit (degree >= linear) already computes per-dimension coefficients internally; this exposes the per-point gradient (dimensions values per point, flattened) in LoessResult.gradient at effectively no extra computation cost. Only supported when surface_mode is "direct" — throws instead of silently leaving gradient as undefined if requested under the default "interpolation" mode. Omitted by default. Gradient values in the overlap region are merged across chunk boundaries the same way y is, via merge_strategy.

See: Intervals

Confidence level for the confidence interval around the mean response (e.g. 0.95), computed per chunk and merged across overlap boundaries the same way y is, via merge_strategy. null (default) disables confidence intervals.

See: Intervals

Confidence level for the prediction interval for new observations (e.g. 0.95); same per-chunk computation and overlap-merging as intervals.confidence. null (default) disables prediction intervals.

Returned by process_chunk() and finalize().

FieldTypeDescription
xFloat64Arrayx values (same order as input)
yFloat64ArraySmoothed y values
fraction_usednumberFraction used
iterations_usednumber | undefinedRobustness iterations actually performed
standard_errorsFloat64Array | undefinedAlways undefined (Batch only)
confidence_lowerFloat64Array | undefinedAlways undefined (Batch only)
confidence_upperFloat64Array | undefinedAlways undefined (Batch only)
prediction_lowerFloat64Array | undefinedAlways undefined (Batch only)
prediction_upperFloat64Array | undefinedAlways undefined (Batch only)
residualsFloat64Array | undefinedResiduals (if "residuals" output)
robustness_weightsFloat64Array | undefinedRobustness weights (if "weights" output)
cv_scoresFloat64Array | undefinedAlways undefined (Batch only)
diagnosticsDiagnostics | undefinedFit metrics (if "diagnostics" output)
gradientFloat64Array | undefinedPer-point local fit gradient, flattened (if "gradient" output, surface_mode = "direct" only)
dimensionsnumberNumber of predictor dimensions
FieldTypeDescription
rmsenumberRoot Mean Squared Error
maenumberMean Absolute Error
r_squarednumberR-squared
residual_sdnumberCumulative sample SD of emitted residuals
effective_dfnumber | undefinedAlways undefined (requires standard errors, Batch only)
aicnumber | undefinedAlways undefined (requires effective_df, Batch only)
aiccnumber | undefinedAlways undefined (requires effective_df, Batch only)