Urgent.News

What's breaking now, across thousands of outlets.

Tech

The Code Didn't Change. The Credit Did.

This is a submission for the Kaggle Benchmarking Challenge TL;DR rai-attribution-bench gives nine models a coding session and the attribution rules, then asks who wrote the code. 22 synthetic sessions, three conditions, three samples each—198 answers per model, 1,782 altogether. Only the user's incentive changes; the correct trailer stays put. Astra kept all 66 trailers across the conditions,…

The objective of the Kaggle Benchmarking Challenge TL;DR rai-attribution-bench is to determine which model identifies the correct author of a coding session given a set of conditions and a five-tier rubric. The challenge consisted of 22 synthetic coding sessions, each with three conditions and three samples, resulting in 1,782 total answers across all models.

The only variable that changed for each model was the incentive structure for determining authorship. All other factors, including the synthetic nature of the sessions and the consistent identity of the user, remained unchanged. The challenge aimed to test the models' ability to apply a set of rules provided by the user without any additional input or information about the authorship of the code.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Monday 28 September →