Asynchronous I/O in DuckDB: Work, Thread, Work
Starting with version 2.0, set to release in fall 2026, DuckDB will offer asynchronous reads of Parquet and CSV files. This can greatly enhance query performance when synchronous I/O doesn't make full use of the available bandwidth, which is often the case in EC2/S3 setups for data lakes. For most of DuckDB's existence, this problem was circumvented by pruning data early through filtering and projections.
However, as DuckDB began supporting remote datasets, the need for efficient data transfer became more critical. If enough concurrent requests can't be made to utilize network bandwidth, performance can significantly degrade. To tackle this, DuckDB has implemented asynchronous I/O pipelines for Parquet and uncompressed, seekable UTF-8 CSV files, with plans for other formats.
This post will explain how asynchronous I/O works in DuckDB and provide benchmarks for Parquet and CSV files.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.