Multimodal AI starts with evidence, not modality
A receiving clerk has to decide whether a refrigerated shipment is safe to accept. Four signals are available: a photograph showing a dented carton and liquid on the floor, a temperature sensor reporting the cold-chain reading, a voice note from the driver describing a hard braking event, the purchase order , which defines the acceptable quantity and packaging. None of these is the truth by…
Every data source available in a given situation contributes evidence, not the overall modality. A receiving clerk must decide if a refrigerated delivery can be accepted. Four sources provide evidence for this decision: a photograph revealing a damaged carton and liquid on the floor, a temperature sensor recording the cold-chain reading, a voice note from the driver detailing a hard braking event, and the purchase order specifying the acceptable quantity and packaging.
None of these sources alone establishes the truth; they serve as evidence produced by a source, captured at a specific time, processed through a pipeline, and interpreted within a policy framework.
A common error is to send all available files to the most powerful model and request a verdict. However, adding more media does not enhance the decision; it introduces irrelevant data, conflicting versions of the same document, unusable information, manipulated content, latency issues, and increased costs. Instead, the decision-making process must begin with the evidence before selecting a single model.
Four stages are involved in this process: sources generate raw evidence, observations extract regions, words, values, and events from this evidence while preserving its origin, alignment joins these observations based on source, position, location, and time, and finally, a bounded decision makes a verdict, abstains, or routes the case to human review. This entire process operates under a consistent policy that remains the same every time.
Beneath this process, an evidence ledger records crucial information for each signal, such as the capture time, transformation details, confidence level of the extraction, applied policy, and a summary of the original data. The key takeaway is that a product should reason over packages of evidence, not just anonymous blobs of media.
For instance, in the shipment case, the system should not only reject but also specify which evidence led to that decision, when each piece of evidence was recorded, and what a reviewer should examine first. This approach ensures that the decision-making process is transparent, reliable, and free from irrelevant or manipulated data.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.