Urgent.News

What's breaking now, across thousands of outlets.

AI

Grew the Collection from 5 Entries to 741 Real Chunks

Every retrieval test since Entry 05 has run against the same handful of entries — three documents in Entries 05-06, twenty after Entry 07's chunking. Realistic enough to prove the mechanics work, not realistic enough to trust the actual rankings. Growing this into something closer to the real published catalog was overdue, and getting there cleanly took two wrong turns first. First attempt…

Since Entry 05, every retrieval test has been conducted against a limited selection of entries. Three documents in Entries 05-06 and twenty in subsequent entries, but these were not realistic enough for reliable rankings. To improve this, the collection needed to grow closer to a real published catalog. Two attempts initially led to issues: the first pointing the bulk-embed script to the wrong folder, resulting in three unrelated files being included, and the second picking up all 45 real files, but re-embedding the 15 "Today I Ran" entries with a different ID scheme.

This led to a mix of chunking styles within the collection, causing self-inflicted version bias. To resolve this, the entire collection was wiped and rebuilt from scratch. The fresh bulk script successfully processed 45 files, resulting in 741 chunks embedded with no duplicates or scaffolding files, and consistent chunking and ID scheme across all documents.

This marked a clean transition from 5 test entries to 741 real chunks drawn from the actual catalog. However, a consideration for future testing is that the collection now includes the entire back catalog of "Today I Ran," potentially causing the original Entry 07 to surface in retrieval tests, affecting the results. The true value of this work lies in having a collection that accurately represents a real knowledge base, enabling more reliable retrieval tests moving forward.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

What the Board Asks Me About AI Agents

I sit in the meetings your architecture decisions eventually reach. Here are the questions that actually get asked about agents at the board level.

  • Board meetings focus on three key concerns about AI agents
  • First concern: blast radius of potential agent mistakes
  • Last concern: references and case studies for agent performance

More from Thursday 8 October →