{
  "id": 10980143,
  "title": "Join Ordering, Part 1: The Shape of the Search Space",
  "url": "https://urgent.news/2026/09/30/join-ordering-part-1-the-shape-of-the-search-space",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-30T11:42:42.000Z",
  "source": {
    "name": "Lobsters",
    "slug": "lobsters",
    "url": "https://deferworks.org/posts/join-ordering/"
  },
  "original_language": "en",
  "account": "In the hidden lore of SQL books, there is no mention of the order in which joins occur, only the desired result they produce. This task falls to the query engine, specifically the planner or optimizer, which aims to select the most efficient join tree. A binary tree structure is used, where the leaves represent base relations and the inner nodes are join operators. While all trees yield the same result for inner joins, their computational costs can vary significantly due to differences in intermediate result sizes.\n\nConsider a scenario with three relations, R₀ and R₁ both containing 1 million rows, and R₂ containing 10 rows. There are two predicates: one between R₀ and R₂, and another between R₂ and R₁. There is no predicate between R₀ and R₁. The first join order, (R₀ ⋈ R₂) ⋈ R₁, quickly reduces R₀ to a few rows and completes efficiently. In contrast, the second order, (R₀ ⋈ R₁) ⋈ R₂, first creates a potentially massive cross-product of up to 10¹² tuples before applying the R₂ predicate, resulting in a ten-fold increase in execution time. This highlights the importance of join ordering in real-world systems, as demonstrated in \"Query optimization through the looking glass,\" where the planner's join choice can lead to execution time variations of 100x or more.\n\nWhile enumerating and evaluating all possible join orders seems like a straightforward solution, the sheer number of possibilities makes this impractical. For example, in the case of three relations, there are 120 distinct join trees, and for five relations, there are 1680 possibilities. The number of join trees rapidly grows, even for a modest 15 relations, reaching approximately 3.5 × 10¹⁸. This problem is NP-hard in general, even when restricted to left-deep trees without cross products. However, many modern query engines can solve moderate instances exactly using dynamic programming techniques.\n\nDynamic programming is based on the simple principle of finding the best plan for all relations in a query. The top join splits the relations into two sets, each of which is recursively planned. This approach allows for the efficient computation of optimal left-deep join trees, even in instances where the computational complexity would otherwise be prohibitive. In the next part of this series, we will delve into the IKKBZ algorithm, which addresses some of the limitations of prior dynamic programming approaches.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}