» NLP · Big Data · 2025

LightningText — Ultra-Lean, Real-Time Text Analytics on Apache Spark

14M+ docs20 KB RAM<2% error8× leaner storage

Built and tuned an Apache Spark pipeline that processes 14 million BookCorpus paragraphs (3.4 GB) in real time. Combined reservoir sampling, Count-Min Sketch, Flajolet–Martin, and wavelet compression to detect heavy hitters and estimate vocabulary with under 2% error using only 20 KB of RAM.

The outcome: 8× leaner storage and minute-level insight — ultra-lightweight NLP suitable for constrained edge environments where traditional pipelines don't fit.