Urgent.News

What's breaking now, across thousands of outlets.

AI

The Qwen3.8-27B Variant Built to Stop Overthinking

Explore Swift-1.5-Qwen3.8-27B-GGUF benchmarks, quantization options, llama.cpp support, limitations, and how it compares with Qwen3.8-27B.

The Qwen3.8-27B Variant Built to Stop Overthinking

Swift-1.5-Qwen3.8-27B-GGUF is a quantized version of the Qwen3.8-27B model, designed to reduce excessive reasoning and speed up processing. UkisAI trained this model to lower the number of thinking tokens by 58.5% on GPQA-Diamond, while increasing its score by 0.31 percentage points compared to the original model. The release also claims a 9.18-fold speed improvement on specific tasks.

The model retains 27 billion parameters, but the exact architecture and hardware requirements are unspecified. To run this model, users must use a llama.cpp-compatible runtime like llama-server.

The key benefit of Swift 1.5 is its lower reasoning-token usage, which is advantageous for coding tasks that require extensive reasoning. In LiveCodeBench v6, Swift 1.5 scored 81.71% compared to 76.76% for the original model, and it used 24.5% fewer reasoning tokens. This makes Swift 1.5 a strong candidate for code generation and problem-solving tasks, provided the outputs can be validated using a compiler or test suite.

The benchmark does not guarantee performance across all programming languages or repository-scale tasks.

Swift 1.5 also showed improvements in tool-using and multi-step agent tasks. On Terminal-Bench 2.1, it scored 72.13% compared to 69.21% for the original model, with mean reasoning tokens decreasing from 52,265 to 43,733. This suggests that Swift 1.5 could be beneficial for tasks that involve a chain of reasoning and tool usage. The model was trained with a focus on long-horizon and agentic tasks.

For general reasoning within a token budget, Swift 1.5 outperformed the base model on GPQA-Diamond, scoring 88.59% compared to 88.28%, while reducing mean tokens from 15,014 to 8,717. This indicates that fewer reasoning tokens can lead to better performance in scenarios where token consumption is a concern. The model is available in different GGUF file sizes, ranging from 8.9 GB to 29.0 GB, allowing users to balance memory usage with performance.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in AI

Twilio Product Chief Warns AI Could Break the Engineering Career Ladder

AI Impact tracks engineering expertise, smarter use of existing power capacity and what failed agents can teach companies.

  • AI could disrupt traditional engineering career ladder by taking on initial coding work.
  • Engineers need to understand system architecture, not just automated code generation.
  • Shani advises focusing on systems engineering and design during onboarding.

More from Friday 2 October →