Why io_uring without readahead is slow: a deep dive into Turso's I/O
A Turso PR implementing readahead for the io_uring backend sparked an investigation into why io_uring with O_DIRECT and no readahead is slower than syscall. The author measures TPC-H Q6, showing that readahead enables request merging, reducing device requests from ~196k to ~16k. They also analyze the cost of SQ polling and cache misses, concluding that O_DIRECT skips the page cache copy, leading to more cache misses during query execution.
With O_DIRECT, though, the disk writes the data into the process buffer directly without the copy step, which means the CPU doesn’t get involved in the process, so nothing is copied to the CPU caches.
- ComputerGuru
Interesting article but it gave me a bit of a panic attack. Benchmarking (with TPC or otherwise) is NOT the way to determine the correct approach here; that is strictly only to be used for databases (typically RDBMS) effectively “owning” the complete hardware they are running on. An embedded database might be used in that manner if it’s operating as the backend for a pure crud application that performs ~zero server side rendering, parsing, validation, etc and is essentially just an async http-to-SQLite interface. But more likely than not, an embedded db will be used and deployed on machines (not necessarily even servers) serving many a purpose, and need to perform best both within the confines of the resources available to the machine and in relative terms, necessarily making tradeoffs that might sacrifice performance for “value” in terms of CPU or memory usage.
This isn’t just with regards to benchmarking, it’s an essential consideration *any* time you are taking ownership of the cache away from the kernel, which is the only piece in the stack that has viability into the global state and can be trusted to give back memory under pressure to ensure everything plays nice together. It’s not limited to just databases or even just memory, for example FreeBSD has had greater than its fair share of issues that trace back to the ZFS having a separate cache from the kernel (despite the much tighter integration between the two and presence of various mechanisms to address pathological […]
- marginalia_nu
If you're doing contiguous readahead in userspace, why not just use preadv? It'll limit you to doing readahead up until the next resident page, but at least in my experiments in Marginalia's index, preadv beats io_uring in all cases you can use a single preadv call to do the full read.
- amluto
I’m curious why the choice is
between syscalls and, specifically, io_uring with O_DIRECT. AFAIK Turso is like SQLite and supports multiple processes accessing the same database, and I would expect buffering to be a huge win in some workloads. What’s wrong with io_uring without direct? There’s also the middle ground of RWF_DONTCACHE.
- ErroneousBosh
This has scrolled past in my feed and every time I've read it as "Io_uring without Radiohead", and I mean you could but what would be the point?
- weatherlite
Does Io_urine use yellow threads?