{
  "id": 12996282,
  "title": "Word-by-word captions with Python and ffmpeg, no CapCut watermark",
  "url": "https://urgent.news/2026/10/09/word-by-word-captions-with-python-and-ffmpeg-no-capcut-watermark",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-09T01:39:29.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/prime619/word-by-word-captions-with-python-and-ffmpeg-no-capcut-watermark-10a"
  },
  "original_language": "en",
  "account": "Word-by-word captions with Python and ffmpeg, no CapCut watermark: An offline script, whisper-word-captions (MIT License), enables the creation of synchronized captions without the need for an account or a watermark. The process involves whisper with word_timestamps set to True, grouping words into short captions (default limit of three words per caption, with additional breaks after punctuation marks or pauses lasting more than 0.6 seconds).\n\nThe Advanced SubStation Alpha (ASS) format is utilized to generate the caption file, with each spoken word assigned a unique Dialogue event. Mid-line style changes can be achieved through curly-brace tags, allowing customization of text color, scaling, and resetting to default settings. A three-word caption example is provided, demonstrating the structure of the caption file.\n\nPython code is used to process the video clip, determine word boundaries, and generate the ASS file. The ffmpeg tool then integrates the generated captions into the video, producing the final captioned video file. Audio remains unaltered, with the video re-encoded once using the libx264 codec at a quality setting of CRF 20. The library used for audio processing is set to copy the original audio.\n\nThe script is compatible with Python 3.10–3.13 and ffmpeg, which must be installed separately. The script downloads the Whisper model during the initial run, taking approximately 490 MB of storage space. The default Whisper model is 'small', optimized for non-English languages. A color preset (highlight colour #FFD400) and thread count (up to 4) are included in the script.\n\nBenchmark results from limited testing indicate that using 8 threads resulted in fewer gaps compared to higher thread counts. The entire process, including transcription and captioning, takes around 6.1–6.7 seconds for a 15-second clip on an 8-core Linux server. A ready-made version of the script is available for purchase on Gumroad, offering additional features such as bundled fonts, color presets, and installation guides.",
  "summary": "Short-form video lives or dies on captions. CapCut and a dozen web tools will do the \"word lights up as you say it\" style for you, usually with a watermark or a free-minute limit. I wanted the same look without an account, so I wrote a small offline script: whisper-word-captions (MIT). This post is about how the trick works, so you can change it. python wordcaptions.py clip.mp4 # ->…",
  "key_points": [
    "Python script creates synchronized captions offline",
    "Uses whisper with wordtimestamps for word boundaries",
    "ffmpeg integrates captions without watermark"
  ],
  "editors_take": "This development lets video creators add synchronized captions without watermarks or accounts, changing how they produce accessible content with open-source tools and customizable styles.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}