Urgent.News

600+ sources. One page. See who else covered it.

Editions

AI

Grok 4.6 Launches With Stronger AI Agents and Frontier-Level Performance

xAI has released Grok 4.6, a new version of its AI model focused on long-running agentic tasks, coding and interactive … Read More The post Grok 4.6 Launches With Stronger AI Agents and Frontier-Level Performance appeared first on ProPakistani .

Grok 4.6 Launches With Stronger AI Agents and Frontier-Level Performance

xAI has unveiled Grok 4.6, an upgraded AI model aimed at tackling intricate, multi-step tasks in areas like coding, visual projects, and interactive development. The new version matches the performance of GPT-5.6 Sol on certain benchmarks and outperforms GPT-5.6 Sol on the Artificial Analysis Intelligence Index. Grok 4.6 is engineered for extended work, excelling at transforming broad ideas into initial functional applications by researching topics, analyzing data, managing codebases, and refining results based on feedback.

It also demonstrates enhanced self-testing and verification during lengthy assignments. Additionally, the model excels in visual and interactive tasks, such as establishing application structures and visual languages in one go, setting a robust foundation for further refinement. Grok 4.6 was trained extensively with additional data focused on reasoning, advanced technical concepts, and specialized environments, including kernel optimization, web development, and computer-aided design.

It underwent a prolonged supplemental training phase and was evaluated on various coding, knowledge-work, and agent benchmarks, achieving top or near-top results across multiple assessments.

Written by urgent.news from ProPakistani's reporting — not their text. Machine-written — it may contain errors, so check the original before relying on it.

Read the original at propakistani.pk →

More in AI

OpenAI Paused Astra for Cyber Risk. Your Agent's Sandbox Escape Is the Same Problem, Smaller Scale

OpenAI paused internal work on its upcoming model, Astra, after evaluations suggested it may have crossed into "Critical" cyber capability territory, including potential autonomous zero-day…

  • OpenAI paused Astra development due to cyber risk concerns.
  • Similar sandbox escape vulnerabilities exploited by Anthropic, Meta, Moonshot.
  • Sentinel proposed to detect unauthorized AI agent behavior and system access.

iris-agentic-dev -- Give Your AI a Live Connection to IRIS, Part 1: The Problem, the Tool, and Getting Started

Part 1 of a series. Part 2 covers the full tool catalog. Part 3 covers ObjectScript skills. Part 4 covers benchmarking and measuring what actually improves.

  • Iris-agentic-dev provides live connection to IRIS for AI assistants
  • Allows AI direct access to entire namespace, including SQL queries and unit tests
  • Open source MCP server built with community, works with major AI tools

One AI Answer, Eight Brands: Designing a Benchmark Without Multiplying the Evidence

I recently ran a small China AI benchmark for eight luxury-jewelry brands. The most interesting result was not a platform ranking. It was a disagreement between two kinds of visibility.

  • Piaget visible in all four answers with China channels, absent from wedding jewelry recommendations
  • Benchmark collected 12 valid raw answers, analyzed separately from raw answers
  • Distinct metrics for different buyer decisions prevent misleading aggregate AI visibility scores