{
  "id": 13511236,
  "title": "Why DuckDB 2.0 is faster",
  "url": "https://urgent.news/2026/10/10/why-duckdb-2-0-is-faster",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-10T18:08:49.000Z",
  "source": {
    "name": "Hacker News",
    "slug": "hacker-news",
    "url": "https://motherduck.com/blog/why-duckdb-20-is-faster/"
  },
  "original_language": "en",
  "account": "DuckDB 2.0, set to launch this fall, has already released an alpha version. A reporter tested its new features on a personal laptop and compared them to the previous version, finding significant performance improvements. The main enhancements include asynchronous I/O, an optimized recursive CTE engine, and VARIANT data type.\n\nAsync I/O is the most noticeable improvement, as it allows network and CPU work to occur simultaneously. This means that instead of waiting for the network to finish each download, the CPU can decode other row groups in the background, resulting in faster data retrieval. The number of downloads running concurrently can be adjusted with the read_ahead_depth setting, which defaults to -1 for automatic adjustment based on the number of threads.\n\nAnother key feature is the rewritten recursive CTE engine, which drastically speeds up queries involving deep parent-child hierarchies, such as git history or bill of materials. Traditional implementations would read the entire table multiple times, once per round, leading to significant slowdowns. In DuckDB 2.0, the table is read only once, and a lookup structure is built to efficiently find the necessary rows for each round. This optimization leads to a 40x speedup for deep hierarchies, making it much more efficient than the previous version.\n\nFinally, DuckDB 2.0 introduces VARIANT, a new data type that intelligently handles JSON data. When writing data to disk, VARIANT looks for consistent fields in JSON columns and extracts them into separate columns for faster querying. Rare or inconsistent fields are stored as a binary remainder alongside the structured data. This approach optimizes storage and CPU usage, especially for large datasets with structured logs. However, for VARIANT to provide its full benefits, the JSON columns should have consistent data types and structures.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}