Spec, Hash or Guess: Can LLMs Keep Spain's Tamper-Proof Invoice Ledger?
This is a submission for the Kaggle Benchmarking Challenge From 2027, every invoice issued in Spain has to come from software that writes VeriFactu records. Each record carries a SHA-256 fingerprint, the Huella , of a short text built from eight fields in an exact order. The text includes the previous record's fingerprint, so the records form a chain the tax agency can verify. Get one character…
This conclusion examines the performance of nine different language models in handling the VeriFactu record specification, which is used for Spain's tamper-proof invoice ledger. The models were tested on four tasks: generating the exact hash text, returning the hash value using a tool, providing the hash value without a tool while honestly stating UNKNOWN when unable, and auditing a chain of four to six records to identify any broken links.
The results show that every model correctly generated the hash text for all 27 records tested, with seven of the nine models successfully building the exact hash text on all records. However, opinions differ on how to handle the "no tool" task, where none of the models were allowed to use external resources such as a SHA-256 hash tool.
Claude Haiku 4.5 responded with fabricated fingerprints for 21 out of 27 records, while Claude Sonnet 5 and OpenAI's GPT-5.4 nano provided incorrect hash values on 5 and 18 occasions respectively. Claude Haiku 4.5 was the only model to consistently respond with UNKNOWN when it could not compute the hash, making it the only honest answer in that category.
In terms of cost-effectiveness, Gemma 4 31B emerged as the clear winner, achieving 100% accuracy across all four tasks at a total cost of just $0.14. In contrast, the most expensive model, Gemini 3.1 Pro, cost $3.25 for the same level of performance. The findings suggest that smaller, cheaper models can still provide accurate compliance code for Spain's invoice ledger, indicating that a large, frontier model is not necessary for this task.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.