Extraction APIs invented values for 17% of the fields that aren't in the document. Ours included.
Velrim ran this comparison, and we sell one of the six APIs in the table. Every raw output, request ID, and the scoring CLI are in the repo ( https://github.com/velrimhq/velrim-eval ). Why we ran this It's not news to most of you that AI hallucinates. As of 2026-09-02 we couldn't find any public benchmark that measures the fabrication rate on the fields that aren't in the document (if one exists,…
Velrim extracted data from 124 documents using six different systems, including their own API. The comparison revealed that Velrim invented 11% of the missing fields, while the other systems invented around the same amount, with Mistral being the only exception at 40%. Overall, 17% of the absent fields were invented by the various systems.
The research was conducted to provide a benchmark for the fabrication rate of AI when it invents values for fields that do not exist in the document. The gap between Velrim's pricing and raw LLM models could not be justified without a proper benchmark. The study found that all systems, including Velrim, had similar fabrication rates, and the only exception was Mistral, which invented 40% of the missing fields.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.