Built and tuned an Apache Spark pipeline that processes 14 million BookCorpus paragraphs (3.4 GB) in real time. Combined reservoir sampling, Count-Min Sketch, Flajolet–Martin, and wavelet compression to detect heavy hitters and estimate vocabulary with under 2% error using only 20 KB of RAM.
The outcome: 8× leaner storage and minute-level insight — ultra-lightweight NLP suitable for constrained edge environments where traditional pipelines don't fit.