StreamingLoess API
See also: fastLoess
When to Use
Section titled “When to Use”- Dataset >100,000 points
- Memory-constrained environments
- Batch processing pipelines
StreamingLoess
Section titled “StreamingLoess”The StreamingLoess class processes data in chunks, suitable for very large datasets or streaming applications.
Constructor:
const { StreamingLoess } = require('fastloess-wasm');
const stream = new StreamingLoess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });console.log("typeof process_chunk:", typeof stream.process_chunk);typeof process_chunk: functionoptions: An object containingStreamingSmoothOptionsfields (a subset of the BatchLoessOptionsfields — see below).streamingOptions: An object containingStreamingOptionsfields.
process_chunk(x, y)
Section titled “process_chunk(x, y)”process_chunk_weighted(x, y, weights) accepts one finite non-negative case weight per observation and preserves weights across buffered overlap points.
Processes a chunk of data. Returns partial results.
const { StreamingLoess } = require('fastloess-wasm');
const n = 100;const x = Float64Array.from({ length: n }, (_, i) => i * 2 * Math.PI / (n - 1));const y = Float64Array.from(x, xi => Math.sin(xi) + 0.1);
const stream = new StreamingLoess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });const partialResult = stream.process_chunk(x.slice(0, 50), y.slice(0, 50));console.log("Fraction used:", partialResult.fraction_used);Fraction used: 0.5finalize()
Section titled “finalize()”Finalizes the smoothing process and returns any remaining buffered results.
const { StreamingLoess } = require('fastloess-wasm');
const n = 100;const x = Float64Array.from({ length: n }, (_, i) => i * 2 * Math.PI / (n - 1));const y = Float64Array.from(x, xi => Math.sin(xi) + 0.1);
const stream = new StreamingLoess({ fraction: 0.5 }, { chunk_size: 50, overlap: 10 });stream.process_chunk(x.slice(0, 50), y.slice(0, 50));stream.process_chunk(x.slice(50), y.slice(50));const finalResult = stream.finalize();console.log("Fraction used:", finalResult.fraction_used);Fraction used: 0.5Options Structures
Section titled “Options Structures”StreamingSmoothOptions
Section titled “StreamingSmoothOptions”| Field | Type | Default | Description |
|---|---|---|---|
fraction | number | 0.67 | Smoothing fraction (bandwidth) |
iterations | number | 3 | Number of robustifying iterations |
weight_function | string | "tricube" | Weight function name |
robustness_method | string | "bisquare" | Robustness method name |
degree | string | "linear" | Polynomial degree of local fit |
dimensions | number | 1 | Number of predictor dimensions |
distance_metric | string | "normalized" | Distance metric; use "minkowski:p" for custom p |
weighted_metric_weights | number[] | null | Per-dimension weights (used when distance_metric = "weighted") |
surface_mode | string | "interpolation" | Surface computation mode |
cell | number | null | Cell size for interpolation grid (smaller → more vertices, higher accuracy) |
interpolation_vertices | number | null | Number of interpolation vertices |
zero_weight_fallback | string | "use_local_mean" | Zero-weight handling strategy |
boundary_policy | string | "extend" | Boundary handling policy |
boundary_degree_fallback | boolean | null | Fall back to lower polynomial degree at boundaries when higher degrees fail |
scaling_method | string | "mad" | Residual scaling method |
auto_converge | number | null | Auto-convergence tolerance |
missing | string | "error" | Policy for non-finite (NaN/Inf) values in each chunk |
parallel | boolean | true | Enable parallel execution |
outputs | string[] | [] | Optional fields: "diagnostics", "residuals", "weights", "gradient" (or "derivative"), "se" |
intervals | { confidence?: number; prediction?: number } | disabled | Grouped confidence and prediction coverage levels. |
"sorted" output and cross-validation are Batch-only and not available here; see fastLoess.
StreamingOptions
Section titled “StreamingOptions”| Field | Type | Default | Description |
|---|---|---|---|
chunk_size | number | 5000 | Data chunk size |
overlap | number | chunk_size / 10 | Overlap between chunks |
merge_strategy | string | "weighted_average" | Strategy for blending overlap regions |
Options
Section titled “Options”fraction
Section titled “fraction”fraction is the most important parameter: it controls the size of the local neighbourhood used at each point.
| Range | Effect | Use case |
|---|---|---|
| 0.1-0.3 | Fine detail | Rapidly changing signals |
| 0.3-0.5 | Balanced | General purpose |
| 0.5-0.7 | Heavy smoothing | Noisy data |
| 0.7-1.0 | Very smooth | Trend extraction |
iterations
Section titled “iterations”iterations controls robustness to outliers, at the cost of speed.
| Value | Effect | Performance |
|---|---|---|
| 0 | No robustness | Fastest |
| 1-3 | Moderate | Recommended |
| 4-6 | Strong | Contaminated data |
| 7+ | Very strong | Heavy outliers |
weight_function
Section titled “weight_function”See: Weight Functions
"tricube"(default)"epanechnikov""gaussian""uniform"(alias:"boxcar")"biweight"(alias:"bisquare")"triangle"(alias:"triangular")"cosine"
robustness_method
Section titled “robustness_method”See: Robustness
"bisquare"(default; alias:"biweight")"huber""talwar"
degree
Section titled “degree”See: Polynomial Degree
"constant"or"0"(degree 0)"linear"or"1"(default, degree 1)"quadratic"or"2"(degree 2)"cubic"or"3"(degree 3)"quartic"or"4"(degree 4)
dimensions
Section titled “dimensions”See: Multivariate LOESS
Number of predictor dimensions. Set to match the number of columns in a multivariate x array.
- Any integer
>= 1;1(default) is univariate
distance_metric
Section titled “distance_metric”See: Multivariate LOESS
"normalized"(default — scales each dimension by its range; alias:"norm")"euclidean"(alias:"euclid")"manhattan"(alias:"l1")"chebyshev"(alias:"linf")"minkowski"(use"minkowski:p"string for custom exponent, e.g."minkowski:3")"weighted"plusweighted_metric_weightsfor per-dimension scaling (alias:"weighted_euclidean")
weighted_metric_weights
Section titled “weighted_metric_weights”See: Multivariate LOESS
Per-dimension weights, one per dimension declared in dimensions. Only used when distance_metric = "weighted"; setting distance_metric = "weighted" without providing this raises an error.
null(default) — has no effect unlessdistance_metric = "weighted"is set- A
number[]of per-dimension weights, required whendistance_metric = "weighted"
surface_mode
Section titled “surface_mode”See: Polynomial Degree
Controls whether the local polynomial is evaluated at every query point or at a sparser grid of anchor vertices with Hermite cubic interpolation in between.
| Mode | Behavior | Speed | Accuracy |
|---|---|---|---|
"interpolation" (default) | Evaluate at vertices, interpolate between | Faster | Slight approximation |
"direct" | Evaluate at every query point | Slower | Full precision |
Cell size for the interpolation grid, as a fraction of the data range. Smaller values place more vertices (denser grid), improving accuracy at the cost of speed. Only applies when surface_mode = "interpolation".
null(default) — uses the library default (0.2)- Any number in
(0, 1]
interpolation_vertices
Section titled “interpolation_vertices”Caps the maximum number of interpolation vertices, overriding the count implied by cell. Only applies when surface_mode = "interpolation".
null(default) — uses the library default (no explicit cap)- Any integer
>= 1
zero_weight_fallback
Section titled “zero_weight_fallback”Behavior when all neighborhood weights are zero:
| Option | Behavior |
|---|---|
"use_local_mean" (default; aliases: "local_mean", "mean") | Use the mean of the neighborhood |
"return_original" (alias: "original") | Return the original y value |
"return_none" (alias: "none") | Return NaN |
boundary_policy
Section titled “boundary_policy”See: Boundary Handling
"extend"(default; alias:"pad")"reflect"(alias:"mirror")"zero""noboundary"(alias:"none")
boundary_degree_fallback
Section titled “boundary_degree_fallback”Whether to reduce the polynomial degree at boundary vertices when the requested degree can’t be fit there (e.g., not enough neighbours). Only applies when surface_mode = "interpolation".
null(default) — uses the library default (enabled)true— falls back to a lower degree at boundariesfalse— raises an error instead of silently falling back
scaling_method
Section titled “scaling_method”See: Scaling Methods
"mad"(default; alias:"median_absolute_deviation")"mar"(alias:"median_absolute_residual")"mean"(alias:"mean_absolute_residual")
auto_converge
Section titled “auto_converge”See: Robustness
Convergence tolerance for early stopping of robustness iterations. null (default) disables early stopping.
missing
Section titled “missing”Policy for handling non-finite (NaN/Inf) values within each chunk:
| Option | Behavior |
|---|---|
"error" (default) | Throw an error if any value in the chunk is non-finite |
"drop" | Silently remove rows where any x dimension or y is non-finite before merging the chunk with the overlap buffer |
Note: A length mismatch between x and y always throws, even under "drop".
chunk_size
Section titled “chunk_size”Number of points processed per call to process_chunk(). Larger chunks reduce per-chunk overhead and give each local fit more surrounding context, at the cost of higher peak memory; smaller chunks bound memory tightly but increase the fraction of points that fall in overlap regions. A good starting point is balancing available memory against how much processing overhead per chunk is acceptable — match it to your file-read buffer or message-batch size to avoid unnecessary copying.
overlap
Section titled “overlap”Number of points retained from the previous chunk as context, so the neighbourhood at chunk boundaries isn’t artificially truncated. Points inside the overlap zone are fitted twice (once by each chunk) and reconciled via merge_strategy. A good starting point is 10–20% of chunk_size: too little overlap causes visible boundary artefacts, while too much wastes computation refitting the same points twice.
null(default) — computeschunk_size / 10, clamped to at least 1 and at mostchunk_size - 10- Any integer
>= 1and< chunk_size
merge_strategy
Section titled “merge_strategy”See: Merge Strategies
| Strategy | Alias | Behavior |
|---|---|---|
"weighted_average" (default) | "weighted" | Distance-weighted blend |
"average" | "mean" | Average overlapping values |
"take_first" | "first" | Keep left chunk values |
"take_last" | "last" | Keep right chunk values |
parallel
Section titled “parallel”Enable multi-threaded execution via the Rayon-based web worker pool.
true(default) — parallelizes the local regression fitsfalse— forces single-threaded execution (useful for benchmarking or deterministic profiling)
outputs: se
Section titled “outputs: se”Include standard errors in the result (LoessResult.standard_errors), computed per chunk and merged across overlap boundaries via merge_strategy.
outputs: diagnostics
Section titled “outputs: diagnostics”Include a Diagnostics object (RMSE, MAE, R², residual_sd) in the result. Streaming residual_sd is the cumulative sample standard deviation of emitted residuals. effective_df/aic/aicc require standard errors, which are Batch-only, so they’re always null here.
outputs: residuals
Section titled “outputs: residuals”Include per-point residuals (y - fitted) in the result.
outputs: weights
Section titled “outputs: weights”Include the final per-point robustness weights (from the last robustness iteration) in the result.
outputs: gradient
Section titled “outputs: gradient”Each local polynomial fit (degree >= linear) already computes per-dimension coefficients internally; this exposes the per-point gradient (dimensions values per point, flattened) in LoessResult.gradient at effectively no extra computation cost. Only supported when surface_mode is "direct" — throws instead of silently leaving gradient as undefined if requested under the default "interpolation" mode. Omitted by default. Gradient values in the overlap region are merged across chunk boundaries the same way y is, via merge_strategy.
intervals.confidence
Section titled “intervals.confidence”See: Intervals
Confidence level for the confidence interval around the mean response (e.g. 0.95), computed per chunk and merged across overlap boundaries the same way y is, via merge_strategy. null (default) disables confidence intervals.
intervals.prediction
Section titled “intervals.prediction”See: Intervals
Confidence level for the prediction interval for new observations (e.g. 0.95); same per-chunk computation and overlap-merging as intervals.confidence. null (default) disables prediction intervals.
Result Structure
Section titled “Result Structure”LoessResult
Section titled “LoessResult”Returned by process_chunk() and finalize().
| Field | Type | Description |
|---|---|---|
x | Float64Array | x values (same order as input) |
y | Float64Array | Smoothed y values |
fraction_used | number | Fraction used |
iterations_used | number | undefined | Robustness iterations actually performed |
standard_errors | Float64Array | undefined | Always undefined (Batch only) |
confidence_lower | Float64Array | undefined | Always undefined (Batch only) |
confidence_upper | Float64Array | undefined | Always undefined (Batch only) |
prediction_lower | Float64Array | undefined | Always undefined (Batch only) |
prediction_upper | Float64Array | undefined | Always undefined (Batch only) |
residuals | Float64Array | undefined | Residuals (if "residuals" output) |
robustness_weights | Float64Array | undefined | Robustness weights (if "weights" output) |
cv_scores | Float64Array | undefined | Always undefined (Batch only) |
diagnostics | Diagnostics | undefined | Fit metrics (if "diagnostics" output) |
gradient | Float64Array | undefined | Per-point local fit gradient, flattened (if "gradient" output, surface_mode = "direct" only) |
dimensions | number | Number of predictor dimensions |
Diagnostics
Section titled “Diagnostics”| Field | Type | Description |
|---|---|---|
rmse | number | Root Mean Squared Error |
mae | number | Mean Absolute Error |
r_squared | number | R-squared |
residual_sd | number | Cumulative sample SD of emitted residuals |
effective_df | number | undefined | Always undefined (requires standard errors, Batch only) |
aic | number | undefined | Always undefined (requires effective_df, Batch only) |
aicc | number | undefined | Always undefined (requires effective_df, Batch only) |