Fast drilldown dashboards from a single Parquet file

Fast drilldown dashboards from a single Parquet file

A single 40MB Parquet data cube, an 18KB reader, an R2 bucket, and a few unassuming http range requests. That's all it takes to serve a fast drilldown dashboard with no database or query engine. The author rolls up 34 million NYC 311 service requests into a cube, stores it on R2, and uses Hyparquet in the browser to fetch only the needed byte ranges. The key is sorting rows so min/max statistics let the reader skip most of the file. The approach shifts complexity to the data pipeline, making it ideal for customer-facing analytics with bounded queries.

In analytics, when all you have is object storage, everything looks like a range request.
  1. simonw

    > The bytes pass through a small Cloudflare Worker on the way, because the free r2.dev URL is rate-limited.

    For a 40MB file I suggest hosting it directly on GitHub Pages - that's effectively a free CORS-enabled CDN and supports HTTP range requests, so you should be able to get that demo working without needing to involve Cloudflare Workers at all.

  2. wyck

    On the surface (I haven't tested it) it looks like a great cost and runtime saver, but only for data sets that need a cadence above 5 or so minutes. You wouldn't be able to have a refresh button to get the "latest" data outside this window, depending on size and build time? Wondering if you can apply a hybrid approach, combining the historical parquet file with a live query.

  3. mrbluecoat

    A clever repurposing of technologies but realistically only worthwhile for static datasets with range payloads small enough to fit into a web response.

    > your pipeline has to rebuild each customer’s file fast enough to meet the update cadence. ... data that updates on a coarse schedule rather than in realtime

More from this day

2026-08-24