Urgent.News

What's breaking now, across thousands of outlets.

AI

Performance of ChatGPT-5.5, Gemini 3, and copilot 2026 in answering chemistry course questions: a cross-sectional study

Scientific Reports, Published online: 10 October 2026; doi:10.1038/s41598-026-73546-z Performance of ChatGPT-5.5, Gemini 3, and copilot 2026 in answering chemistry course questions: a cross-sectional study

We haven't written up this one. Scientific Reports has the full story — the link below goes straight to it.

Read the original at nature.com →

More in AI

What Is Physical AI? Nvidia Is the Stock I'd Buy to Own It.

A rival chipmaker agreed last month to pay $8.2 billion in shares to push further into this area of AI.

  • Nvidia uses "Physical AI" for AI operating in the real world
  • Nvidia's automotive revenue of $2.35 billion is 1% of total $215.9 billion revenue
  • Nvidia's physical AI segment is less than 3% of total sales at $6 billion

Deterministic State Machines for Resilient Autonomous Agents

Deterministic State Machines for Resilient Autonomous Agents Autonomous multi-agent architectures routinely fail in production when relying on unconstrained large language model conversation loops.

  • Deterministic FSM replaces unconstrained LLM loops to prevent unpredictable behavior
  • AgentStateGraph governs deterministic control flow outside LLM reasoning core
  • Strict JSON validation and checkpointing enable state persistence and recovery

We quantized our AI judge. Here's exactly what broke.

Our production judge — a small 1.7B model with a LoRA adapter that grades other AI outputs as pass / fail / insufficient_evidence (88.5% accuracy, ECE 0.072) — is cheap to run.

  • Quantization implemented to reduce serving costs
  • Precision dropped to 98.28% and 94.16% with int8 and int4 formats
  • Four-rule deployment discipline established for production judges

More from Saturday 10 October →