{
  "id": 10394124,
  "title": "Scraping YouTube transcripts at scale for RAG (and the 3 things that break it)",
  "url": "https://urgent.news/2026/09/28/scraping-youtube-transcripts-at-scale-for-rag-and-the-3-things-that",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-28T07:23:36.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/casaucao/scraping-youtube-transcripts-at-scale-for-rag-and-the-3-things-that-break-it-2hdf"
  },
  "original_language": "en",
  "account": "If you are constructing a RAG or agent pipeline for video content, the transcript serves as the primary data source. While titles and descriptions are limited, a 40-minute video contains thousands of tokens of dense, quotable text. The good news is that YouTube provides this text through its captions. The bad news is that most amateur pipelines fail when attempting to extract captions for videos they do not own at scale. This article outlines the authoritative source of truth, the useless official API, a functional Python method for loading the data into a vector store, and the pitfalls to design around.\n\nCaptions are preferred over Automatic Speech Recognition (ASR) for several reasons. While ASR can work for single videos, it becomes costly, slow, and introduces a second error source since the ASR output may not match the captions that viewers already see. YouTube captions are synchronized to the media, carry segment timestamps, and are often human-authored for important videos. When captions are absent, auto-generated captions are still a superior starting point compared to attempting to transcribe the audio yourself.\n\nHowever, it is important to note that caption extraction is not transcription. If a video lacks captions altogether, there is nothing to extract. In such cases, a real ASR step would be necessary, and this scenario should be handled explicitly rather than resulting in an empty document. The official YouTube Data API v3 cannot fetch captions from third-party videos. Its captions.download endpoint only works for videos owned by the authenticated channel or those for which you have permission. For arbitrary third-party videos needed by your RAG corpus, the API returns an error. The correct route is leveraging the public player surface: the watch page's player response contains the caption track metadata, which points to timedtext resources that can be fetched without requiring OAuth or a Google Cloud key. This approach mirrors the method used by YouTube's own web player and does not necessitate additional authentication or API access.\n\nTo simplify the process, you can utilize an Actor to handle the player payload, manage retries, and implement proxy rotation. Here is the complete workflow:\n\n1. Submit a batch of video URLs.\n2. Retrieve the dataset containing the transcripts.\n3. Convert each transcript into chunk records suitable for embedding.\n\n```python\nfrom apify_client import ApifyClient\n\nclient = ApifyClient(APIFY_TOKEN)\n\nrun = client.actor(casaucao/youtube-transcript-scraper).call(\nrun_input={\n\"videoUrls\": [\n\"https://www.youtube.com/watch?v=dQw4w9WgXcQ\",\n\"https://youtu.be/9bZkp7q19f0\",\n],\n\"languages\": [\"en\", \"es\"],\n\"outputFormat\": \"segments\",\n\"includeMetadata\": True,\n\"maxRetries\": 3,\n\"concurrency\": 5,\n}\n)\n\nchunks = []\nfor item in client.dataset(run[defaultDatasetId]).iterate_items():\nif item.get(\"status\") != \"ok\" or not item.get(\"fullText\"):\ncontinue\n\nsegments = item[\"segments\"]\nwindow = []\nsize = 0\nstart = segments[0][\"start\"] if segments else None\n\nfor seg in segments:\nwindow.append(seg[\"text\"])\nsize += len(seg[\"text\"])\n\nif size >= 1200:\nchunks.append(\n{\n\"videoId\": item[\"videoId\"],\n\"title\": item[\"title\"],\n\"language\": item[\"language\"],\n\"start\": start,\n\"text\": \" \".join(window),\n}\n)\nwindow, size = [], 0\nstart = seg[\"end\"]\n\nif window:\nchunks.append(\n{\n\"videoId\": item[\"videoId\"],\n\"title\": item[\"title\"],\n\"language\": item[\"language\"],\n\"start\": start,\n\"text\": \" \".join(window),\n}\n)\n\nprint(len(chunks), \"chunks ready to embed\")\n```\n\nThe key fields to focus on are fullText for the entire document, segments for maintaining timestamp consistency during chunking, language and translatedFrom for provenance information, and isGenerated to determine whether the text was written by a human or generated by a machine. After obtaining the chunks, feed each text into your embedding model and store the videoId and title as metadata to enable proper citations linking back to the original source.",
  "summary": "If you are building a RAG or agent pipeline over video, the transcript is the payload. Titles and descriptions are thin; a 40-minute talk is thousands of tokens of dense, quotable prose. The good news is that YouTube already exposes that prose as captions. The bad news is that pulling captions for videos you do not own, at batch scale, is where most naive pipelines fall over. This post walks…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}