Fast drilldown dashboards from a single Parquet file

A single 40MB Parquet data cube, an 18KB reader, an R2 bucket, and a few unassuming http range requests. That's all it takes to serve a fast drilldown dashboard with no database or query engine. The author rolls up 34 million NYC 311 service requests into a cube, stores it on R2, and uses Hyparquet in the browser to fetch only the needed byte ranges. The key is sorting rows so min/max statistics let the reader skip most of the file. The approach shifts complexity to the data pipeline, making it ideal for customer-facing analytics with bounded queries.
In analytics, when all you have is object storage, everything looks like a range request.
- simonw
> The bytes pass through a small Cloudflare Worker on the way, because the free r2.dev URL is rate-limited.
For a 40MB file I suggest hosting it directly on GitHub Pages - that's effectively a free CORS-enabled CDN and supports HTTP range requests, so you should be able to get that demo working without needing to involve Cloudflare Workers at all.
- wyck
On the surface (I haven't tested it) it looks like a great cost and runtime saver, but only for data sets that need a cadence above 5 or so minutes. You wouldn't be able to have a refresh button to get the "latest" data outside this window, depending on size and build time? Wondering if you can apply a hybrid approach, combining the historical parquet file with a live query.
- mrbluecoat
A clever repurposing of technologies but realistically only worthwhile for static datasets with range payloads small enough to fit into a web response.
> your pipeline has to rebuild each customer’s file fast enough to meet the update cadence. ... data that updates on a coarse schedule rather than in realtime