The 0.87 problem: when semantic linking makes inconsistent records look connected
A 0.87 match that almost shipped Last quarter I built a small semantic-similarity layer over our CAPA database. The pitch to myself was simple: an engineer files a CAPA, the system surfaces related complaints, prior CAPAs, and design changes with confidence scores. The faster the linkage, the faster the impact assessment. The faster the assessment, the more robust the closure. It worked. Too…
In the world of quality management for medical devices, a semantic similarity layer was built over a CAPA database to help engineers find related complaints, prior CAPAs, and design changes with confidence scores. While the model performed well, it also produced inaccurate matches that could lead to flawed decisions. For instance, a CAPA related to a supplier-side dimensional drift on a machined titanium housing was flagged as 0.87 semantically similar to a design change from several months prior, despite being from a different supplier, alloy, and revision.
The issue lies in the inflation problem, where inconsistent records can appear connected due to the semantic linking. This can create a false sense of security and lead to a decision being made based on a similarity score rather than a solid decision chain. The ISO 13485:2016 standard, clause 8.5.2(f), requires a review of the effectiveness of any corrective action taken, but the similarity score itself is not a decision, but rather a hint dressed up as one.
To address this issue, the following three steps were introduced after the incident: source consistency check, explicit human ratification, and traceable decision lineage. Before citing any related record in a CAPA, the underlying data sources must be reconciled and any discrepancies resolved. The CAPA owner must write a plain language explanation as to why one record is related to another, rather than simply stating that the model says the records are related.
Lastly, the decision and not the similarity score should be retrievable from the CAPA file, allowing the auditor to review the evidence without needing the model to run.
While semantic linking can be a helpful tool, it is only useful when it serves as assistance and not authority. The QA lead should keep the score at the front of the investigation and not the back of the closure. Ultimately, the open question remains: if a semantic-similarity layer is the only basis for relating two records, and an auditor asks the CAPA owner to justify the linkage from memory and the file, what evidence would be present today to support their response?
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.