ProofSec: Benchmarking Epistemic Robustness and Evidence-Grounded Vulnerability Reasoning in Frontier LLM
This is a submission for the Kaggle Benchmarking Challenge What happens when an LLM recognizes every lexical and semantic signature associated with a vulnerability - IDOR, BOLA, authorization bypass, predictable identifiers - but the available evidence does not actually establish that the vulnerability exists? That question is the foundation of ProofSec , an evidence-centric security reasoning…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.