Urgent.News

What's breaking now, across thousands of outlets.

AI

Twitter.now trust scores face test on disputed claims

Twitter. now is betting that artificial intelligence can help users screen misinformation, but research on automated fact-checking suggests the technology becomes substantially less dependable when claims are politically contested, context-heavy or difficult to verify. The social network, developed by Virginia-based Operation Bluebird, has placed a system called VERA at the centre of its effort…

Twitter is banking on artificial intelligence to combat misinformation, but research indicates that automated fact-checking becomes less reliable when dealing with politically charged claims, complex contexts, or claims that are difficult to verify. The social platform, created by Virginia-based Operation Bluebird, has introduced a system called VERA, which is the core of their strategy to differentiate themselves from X. VERA examines posts, assigns trust signals, and works alongside a user-controlled Trust Dial, which can reduce the visibility of content deemed untrustworthy.

Some media coverage has portrayed VERA as utilizing Google's Gemini artificial intelligence models to evaluate factual claims, provide context, and generate numerical trust scores. However, the company's website still describes VERA and its broader Trust OS as "coming soon," suggesting a gap between prototype testing and full deployment across the network.

The central challenge for automated verification lies in accurately assessing claims that involve changing information, conflicting evidence, political interpretation, or cultural context. Studies have shown that large language models perform inconsistently in these situations. While straightforward cases with easily confirmed facts can be handled accurately and swiftly, more complex cases often require human oversight, explanation, and participation rather than just confidence scores.

Other research has highlighted similar shortcomings in automated fact-checking. An experimental study of political headlines found that an AI system correctly identified 90% of the false headlines it examined, but its performance on true material was significantly lower. Only three out of 20 true headlines were accurately identified as true, four were mistakenly labeled as false, and the remaining 13 provided uncertain answers.

Human-written fact checks have proven more effective in helping people distinguish accurate from inaccurate information.

Furthermore, research published in Nature Machine Intelligence has demonstrated that language models struggle to differentiate between knowledge, belief, and factual propositions, depending on how information is framed and whose beliefs are being represented. Tests involving 24 models and approximately 13,000 questions revealed significant performance differences based on information framing and represented beliefs, indicating limitations in treating AI-generated confidence assessments as equivalent to factual certainty.

Despite these challenges, automated systems can still be beneficial when used cautiously. Research on automated climate fact-checking has shown that advanced language models, combined with structured evidence, can evaluate specialized claims. Studies on misinformation detection have also revealed that AI can help process large volumes of material that would otherwise overwhelm human fact-checkers.

Field evaluations on X have produced promising results, with an AI system generating Community Notes writing over 1,600 notes and receiving favorable helpfulness ratings from users of various political leanings.

However, Twitter's approach raises new questions about transparency. A trust score can influence how extensively a claim is disseminated, even if the original post remains online. This could potentially reduce the visibility of legitimate information, particularly in cases involving satire, developing news, disputed political statements, or poorly represented languages.

The performance of these models can vary based on factors such as language, subject matter, source quality, and the contextual evidence provided to the system. This underscores the importance of considering the limitations and biases inherent in automated fact-checking technologies.

Written by urgent.news from Arabian Post's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thearabianpost.com →

More in AI

Building an AI Forensic Investigator for Vehicle Failures

I built an AI agent that diagnoses cars — and asks before it touches anything Built for the TrueForge Agent Harness Hackathon (Aug 24–30, 2026). "Something expensive broke.

  • FaultTrace AI system developed for vehicle failure diagnostics
  • Uses sensor data, creates hypotheses, runs sandbox analysis
  • Stops for human approval before any physical actions

More from Saturday 29 August →