Urgent.News

What's breaking now, across thousands of outlets.

AI

Keeping One Character Consistent Across a Whole Article's AI Illustrations

If you've tried to illustrate a long blog post with an image model, you know how it goes. You ask for five pictures and get five different art styles, three versions of your "mascot", and a stock-photo handshake you never asked for. Each image is fine by itself. Put them together in one post, though, and it looks like a ransom note. I build InkDoo , a tool that takes an article and returns a full…

Creating a series of consistent illustrations for a lengthy blog post using an image generation model can be challenging. InkDoo is a tool that addresses this issue by crafting a full set of hand-drawn explainer illustrations that share the same character throughout the article. Consistency is achieved through several key aspects in the process.

First, each character is defined as a visual spec, providing exact instructions and specifications. For instance, the default character, Mochi, is described as a round, low, squishy cat-like blob drawn as a hollow black outline with specific details such as two tiny triangle ears, two short sleepy horizontal line eyes, a tiny 'w' mouth, small nub paws, and one thin curly tail. By describing the character in this manner, the model receives precise instructions, reducing the chances of inconsistency.

Second, the character is made to perform the core conceptual action of each image. This involves writing the scene description and ensuring that the character is involved and visible in every image. By doing this, the character becomes an essential part of the metaphor, ensuring that it remains recognizable and consistent across all illustrations.

Third, the visual style is maintained separately from the content. Each image prompt has a fixed visual DNA block, which includes elements such as a pure white background, minimalist black hand-drawn wobbly line art, a certain level of empty space, and sparse handwritten notes in specific colors. This approach ensures that the style remains consistent across all illustrations while allowing for subtle variations in theme, structure type, core idea, and composition.

Fourth, the image generation model does not directly receive images uploaded by users. Instead, it relies on a text twin provided by the user. This enables the planner LLM to work with the character's specifications without having access to the image itself. The planner LLM takes the user's character description and creates a text-only version to be used by the image model, ensuring a consistent look throughout the series.

Fifth, switching characters without re-planning is possible in InkDoo. Since the scenes are written in English and contain the character's name, switching between characters merely involves a careful, Unicode-aware whole-word replace of the old name with the new one. This preserves consistency without requiring a new planning session.

Finally, the exported Markdown (or WeChat-friendly HTML) contains the hosted image URLs at the appropriate locations within the article. The planner returns an after anchor for each shot, representing the first 20 characters of the paragraph where the image should follow. By matching on a short prefix, the exporter can accurately place each image, avoiding issues that would arise from asking the model for paragraph numbers, which may lead to inaccuracies.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

The AI buildout has a power problem

The first piece in this series argued the AI buildout is a financing story: $700-730 billion in 2026 hyperscaler capex, up 70-80% from last year, and money is not the constraint.

  • AI buildout requires significant power expansion.
  • US data centers could consume 649 TWh by 2030.
  • Major tech firms contract for new nuclear power.

More from Tuesday 6 October →