Skip to content

Genomic Data Smoothing

LOWESS for methylation profiles, ChIP-seq signals, and other genomic data.

Genomic data often contains noise from sequencing depth variation, PCR artifacts, or biological heterogeneity. LOWESS smoothing helps reveal underlying patterns.


DNA methylation data (from bisulfite sequencing or arrays) shows position-dependent patterns that can be obscured by measurement noise.

A small fraction = 0.1 lets LOWESS follow fine-scale spatial structure without smearing the transitions between methylated and unmethylated regions. confidence_intervals = 0.95 produces uncertainty bands that naturally widen at positions with sparser CpG coverage, making low-confidence segments immediately apparent in the plot.

const fl = require('fastlowess');
const n = 100;
const x = Float64Array.from({ length: n }, (_, i) => i * 2 * Math.PI / (n - 1));
const y = Float64Array.from(x, (xi, i) => Math.sin(xi) + (((i*7+3)%17)/17-0.5)*0.6);
const positions = Float64Array.from({ length: 1000 }, (_, i) => i * 10.0);
const observed = Float64Array.from(positions, (p, i) => 50 + Math.sin(p/100)*20 + ((i*7+3)%17)/17*5);
// positions and observed are your methylation data (Float64Array)
const model = new fl.Lowess({
fraction: 0.1,
iterations: 3,
confidence_intervals: 0.95
});
const result = model.fit(positions, observed);
// Smoothed profile in result.y
// CI bounds in result.confidence_lower/upper
console.log("95% CI: [" + result.confidence_lower[0].toFixed(4) + ", " + result.confidence_upper[0].toFixed(4) + "]");
95% CI: [54.5520, 58.9813]

ChIP-seq experiments produce sparse, noisy coverage data. LOWESS can help identify binding regions.

fraction = 0.05 provides high spatial resolution — important for resolving narrow binding peaks that would otherwise be smeared into the background. The larger iterations = 5 is deliberate: Poisson-distributed read counts produce tall, isolated spikes, and extra robustness iterations progressively down-weight them so the estimated background level is not inflated by a handful of extreme counts.

const fl = require('fastlowess');
const n = 100;
const x = Float64Array.from({ length: n }, (_, i) => i * 2 * Math.PI / (n - 1));
const y = Float64Array.from(x, (xi, i) => Math.sin(xi) + (((i*7+3)%17)/17-0.5)*0.6);
const positions = Float64Array.from({ length: 1000 }, (_, i) => i * 10.0);
const observed = Float64Array.from(positions, (p, i) => 50 + Math.sin(p/100)*20 + ((i*7+3)%17)/17*5);
const model = new fl.Lowess({
fraction: 0.05,
iterations: 5
});
const result = model.fit(positions, observed);
// Identify peaks above threshold
const smoothed = result.y;
const threshold = 50.0; // Example threshold
const peaks = positions.filter((p, i) => smoothed[i] > threshold);
console.log("Peak count:", peaks.length, " smoothed[0]:", result.y[0].toFixed(2));
Peak count: 570 smoothed[0]: 58.00

For whole-genome data that doesn’t fit in memory:

const { StreamingLowess } = require('fastlowess');
const positions = Float64Array.from({ length: 1000 }, (_, i) => i * 10.0);
const observed = Float64Array.from(positions, p => 50 + Math.sin(p/100)*20 + Math.random()*5);
// Array of genomic chunks to process
const genomicData = [
{ positions: positions.slice(0, 500), coverage: observed.slice(0, 500) },
{ positions: positions.slice(500), coverage: observed.slice(500) }
];
const processor = new StreamingLowess(
{ fraction: 0.05, iterations: 3 },
{ chunk_size: 100000, overlap: 10000 }
);
// Process genomic chunks from stream or file
for (const chunk of genomicData) {
processor.process_chunk(chunk.positions, chunk.coverage);
}
const result = processor.finalize();
console.log("Smoothed", result.y.length, "points via streaming");
Smoothed 1000 points via streaming

ConsiderationRecommendation
Fraction0.05–0.15 (preserve local features)
Iterations3–5 (handle sequencing outliers)
Large dataUse streaming mode
Sparse regionsUse boundary_policy="extend"
Multiple chromosomesProcess separately or ensure sorted