Skip to content

Benchmarks

These shared native CPU benchmark results compare base R and serial/parallel rfastlowess; they do not measure WebAssembly, whose binding is single-threaded. Timings depend on hardware, thread availability, and workload, so results are reference observations rather than universal thresholds.

Runtime and speedup comparison of stats::lowess with serial and parallel rfastlowess across benchmark categories

The plot and table use mean CPU timings. Parentheses show speedup relative to stats::lowess; values above 1 indicate faster execution than R.

Scenariostats::lowessrfastlowess (serial)rfastlowess (parallel)
clustered2.34 ms2.15 ms (1.1×)1.07 ms (2.2×)
constant_y1.62 ms2.12 ms (0.8×)0.76 ms (2.1×)
extreme_outliers6.34 ms6.35 ms (1.0×)2.91 ms (2.2×)
financial_10000.26 ms0.22 ms (1.2×)0.22 ms (1.2×)
financial_5000.22 ms0.14 ms (1.6×)0.15 ms (1.5×)
financial_50001.12 ms0.92 ms (1.2×)0.79 ms (1.4×)
fraction_0.050.97 ms0.76 ms (1.3×)0.93 ms (1.1×)
fraction_0.11.67 ms1.39 ms (1.2×)1.05 ms (1.6×)
fraction_0.22.58 ms2.28 ms (1.1×)1.20 ms (2.1×)
fraction_0.33.65 ms3.20 ms (1.1×)1.28 ms (2.9×)
fraction_0.54.97 ms4.98 ms (1.0×)1.87 ms (2.7×)
fraction_0.676.07 ms6.73 ms (0.9×)1.84 ms (3.3×)
genomic_10000.23 ms0.28 ms (0.8×)0.34 ms (0.7×)
genomic_10000032.71 ms25.84 ms (1.3×)8.76 ms (3.7×)
genomic_50001.42 ms1.57 ms (0.9×)1.10 ms (1.3×)
high_noise8.01 ms7.46 ms (1.1×)2.33 ms (3.4×)
iterations_00.48 ms0.50 ms (1.0×)0.36 ms (1.3×)
iterations_11.31 ms1.47 ms (0.9×)0.60 ms (2.2×)
iterations_107.29 ms5.96 ms (1.2×)2.63 ms (2.8×)
iterations_21.95 ms2.09 ms (0.9×)1.07 ms (1.8×)
iterations_32.57 ms2.19 ms (1.2×)1.10 ms (2.3×)
iterations_53.37 ms3.33 ms (1.0×)1.62 ms (2.1×)
large_delta_010 067.77 ms5 279.61 ms (1.9×)1 434.48 ms (7.0×)
large_delta_0.111.90 ms12.68 ms (0.9×)5.05 ms (2.4×)
large_high_fraction9 434.85 ms4 940.71 ms (1.9×)1 216.33 ms (7.8×)
large_high_iter31 189.44 ms14 409.17 ms (2.2×)4 475.26 ms (7.0×)
scale_10000.38 ms0.36 ms (1.1×)0.35 ms (1.1×)
scale_100002.69 ms2.59 ms (1.0×)1.51 ms (1.8×)
scale_50001.82 ms1.43 ms (1.3×)0.99 ms (1.8×)
scientific_10000.42 ms0.60 ms (0.7×)0.50 ms (0.8×)
scientific_5000.29 ms0.25 ms (1.2×)0.30 ms (1.0×)
scientific_50001.92 ms1.88 ms (1.0×)1.35 ms (1.4×)

At small sizes, fixed per-call overhead can dominate, and parallel execution is not always faster. The large wide-window workload reaches 7.8× parallel speedup; the high-iteration workload reaches 2.2× serial speedup. Allowing default delta reduces the reference workload from approximately 10.07 s to 11.90 ms.