Why DuckDB 2.0 is faster
Article URL: https://motherduck.com/blog/why-duckdb-20-is-faster/ Comments URL: https://news.ycombinator.com/item?id=50035530 Points: 200 # Comments: 62
DuckDB 2.0, the upcoming release, promises significant performance improvements. In an alpha test, the author compared the new version against the previous one and found faster speeds, especially when dealing with large data sets on S3 storage. The key performance enhancements revolve around three features: async I/O, rewritten recursive CTE engine, and VARIANT data type.
Firstly, async I/O significantly speeds up data reading from S3 storage. In the test, DuckDB 2.0 was able to download and decode 2268 row groups, each containing about 122,000 rows, while crunching on the CPU simultaneously. This parallel processing eliminated waiting time, doubling or tripling the speed of data retrieval compared to version 1.5.5. This improvement is automatic and does not require any changes to existing queries; just set the read_ahead_depth to -1.
Secondly, DuckDB 2.0 introduces a rewritten recursive CTE engine. This feature is particularly useful for handling complex data structures such as an org chart or a git history, where data is organized in a parent-child relationship. The traditional approach in version 1.5.5 involved re-reading the entire table for every round, creating a linear increase in read operations.
In contrast, the new engine reads the table once, builds a lookup table for the parent column, and then employs efficient lookups for each round. Consequently, this leads to linear growth in execution time as the hierarchy depth increases. For instance, walking through a 20,000-commit git history now becomes a feasible operation on a standard laptop.
Lastly, the introduction of the VARIANT data type provides better handling of JSON data. By analyzing the most frequently occurring fields within a JSON column, DuckDB 2.0 can extract those consistent fields into separate columns while preserving the irregular data in a binary remainder. This technique, known as shredding, optimizes storage and processing for structured logs or other JSON-heavy datasets.
However, it's essential to maintain consistency in the value types of the consistent fields to avoid slower processing due to type conversion.
Overall, DuckDB 2.0 offers notable performance gains for specific use cases, enabling organizations to leverage powerful data analysis capabilities on their existing infrastructure without requiring significant modifications to existing queries or data structures.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- Why DuckDB 2.0 is faster motherduck.com