Your seed data is lying to you — the case for foreign-key-consistent test data
Here's a bug I've shipped more than once: tests pass locally, demo looks great, then something breaks the moment real relational data shows up. The culprit is almost always the same — seed data with foreign keys that don't actually point at anything. The quiet problem with hand-rolled seed data Say you've got the classic Prisma setup: model User { id Int @id @default(autoincrement()) email String…
I've encountered a recurring issue in my development process: tests passing locally but breaking when real relational data is introduced. The root cause is often seed data with foreign keys that fail to point to existing records. This common problem arises when using hand-rolled seed data, especially in a Prisma setup.
When you generate authorId as a random integer using a generic faker-style approach, the data appears fine in a JSON preview. However, a JOIN between Post and User may return rows that shouldn't exist or show none at all. A real foreign key constraint in Postgres will reject the insert, violating the foreign key constraint. Any UI rendering posts by author will either crash or display ghost data. While your seed data may appear valid at first glance, it ultimately fails the crucial test of behaving like real data.
The issue lies in the fact that faker (and similar tools) excel at generating values like names, emails, and prices, but they lack understanding of relationships. They can easily assign authorId as 8471 to a post when no user with that id exists, because each field is generated in isolation. The generated values are realistic, but the relationships are fictitious.
To generate authentic relational data, you need to follow these steps: generate parent rows first (User), remember their actual ids, generate child rows (Post), and set each authorId to an id that was actually generated for a user. While this concept is straightforward, it can be tedious to implement manually each time you modify the schema, which is why people often skip this step and ship dangling keys. Automating this process can save time and ensure data integrity.
I developed a VS Code extension called SeedForge that reads your schema.prisma and generates data for you. The free edition handles single-table generation with realistic typed values and exports to JSON, SQL, or CSV. The core functionality of generating related tables, so the child's foreign key column points to a real generated parent, is what took the most effort to get right. This feature is available on Open VSX if you're using Cursor or VSCodium.
Regardless of whether you use SeedForge or not, it's essential to ask yourself after writing seed data whether your foreign keys reference rows that exist. If they don't, your tests are essentially rehearsing against a scenario that can't occur in production. Generate parent records first and reuse their real ids for children to avoid these issues.
Your JOINs will then accurately reflect the truth. What's your approach to relational test data - hand-written factories, SQL fixtures, or something else? I'm curious to learn about the methods people find effective.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.