Urgent.News

What's breaking now, across thousands of outlets.

AI

A Scientist Was Presenting His Vibe-Coded New Work When One of His Colleagues Pointed Out Something Extremely Embarrassing

"It was subtle, but it was very wrong, and everything downstream of it was also wrong, and I had shared the whole thing in a room full of people who trusted me." The post A Scientist Was Presenting His Vibe-Coded New Work When One of His Colleagues Pointed Out Something Extremely Embarrassing appeared first on Futurism .

A Scientist Was Presenting His Vibe-Coded New Work When One of His Colleagues Pointed Out Something Extremely Embarrassing

Renowned astrophysicist and science communicator Paul Sutter shared a cautionary tale in a recent article for Nautilus. During a presentation in February to a group of collaborators, Sutter unveiled a new algorithm he had created to identify voids, or empty spaces between galaxies. The update boasted impressive improvements, being ten times faster, employing a more nuanced approach to data handling, and capable of processing surveys a hundredfold larger than before.

Sutter had even utilized an AI tool to assist in writing the code, a method he referred to as "vibe coding."

However, just ten minutes into his talk, a colleague raised an alarm, expressing that something appeared amiss. Upon further examination, Sutter admitted that the new algorithm had been mishandling the edges of the surveys, leading to incorrect results that cascaded throughout his presentation. It wasn't a simple typo or missing citation, but rather a subtle yet significant error. Sutter had shared the entire presentation with a group of people who trusted his expertise, putting them all at risk.

Sutter's story highlights the pervasive issue of AI-generated content that may appear authoritative and convincing but is, in fact, riddled with inaccuracies. In the academic world, where research must be rigorous and well-supported, this poses a serious concern. The prevalence of poorly-researched and often unedited AI-generated content has led to a need for academics to be vigilant and responsible for every piece they publish, including any embarrassing AI-induced errors.

Sutter likens the current state of AI use to the era of alchemy, an ancient practice that lacked scientific rigor. He asserts that we are currently in the pre-chemistry era of AI, where users, like the alchemists of old, are not stopping their reliance on the tool. To navigate this new landscape, Sutter recommends carefully tracing an AI's chain of reasoning and thoroughly auditing its outputs. He vowed to work differently moving forward, making a conscious effort to distrust AI-generated information.

Interestingly, when Sutter's latest piece was analyzed using an AI detection tool, it was found that 56% of the text seemed to have been generated by an AI. This finding underscores the importance of being aware of the risks associated with AI-generated content, even for the most accomplished thinkers.

Written by urgent.news from Futurism's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at futurism.com →

More in AI

The #1 row on this AI memory leaderboard is not a measurement

Bench'd (benchd.ai) calls itself the neutral benchmark authority for AI memory, and sells vendors a verification badge from $299 to $3,999.99 a month. I ran my memory system through their harness.

  • AI memory leaderboard's top entry is not genuine measurement
  • Benchmark service Bench d sells verification badge for $299-$3,999.99/month
  • Top three entries lack downloadable proof of Community-Verification

I built an AI incident responder that refuses to fix anything without asking

Built for the WeMakeDevs × TrueFoundry Agent Harness Hackathon. There are two kinds of "AI for incident response," and both of them are wrong. The first acts on its own.

  • Mayday AI incident responder combines speed of autonomous action with safety of human approval
  • Proposes single fix with root cause, impact, and ruled-out options before human approval
  • 401 HTTP error prevents bypassing approval process

More from Saturday 29 August →