Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
The latest iteration of open models, Qwen 3.8, has been evaluated in a reasoning prefills experiment, building upon the work conducted with GPT-5.5 Pro. The study compares the two models by measuring the proportion of the teacher model's visible answer that appears in the first 100 tokens of the target model's response. The results reveal that Qwen has made significant progress, achieving a +20.58 point increase in alignment with GPT-5.5 Pro compared to its previous performance.
This improvement is particularly notable in the private synthetic puzzles, suggesting that Qwen may have learned from GPT-5.5 Pro or a closely related GPT model instead of Opus. Furthermore, Kimi K3 also demonstrates a high overlap with GPT-5.5 Pro, with values of 50.11% without the prefills and 54.42% with the prefills. However, the prefills only contribute an additional +4.31 points to the alignment score.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.