Urgent.News

What's breaking now, across thousands of outlets.

AI

My weather app said the sky would be clear. I taught an open model to check, then went outside

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Every stargazer knows this evening. The app says clear . You find the warm jacket, drive out past the streetlights, let your eyes adjust for twenty minutes… and look up at a flat grey lid. Cloud cover is the one number that decides whether going outside at night is worth it, and it's the number weather apps…

Cloud cover is the deciding factor for whether going outside at night is worth it, and weather apps often get this wrong. The author, a systems engineer, decided to build a solution to address this issue. He created Clear Tonight, an open-source tool that compares the cloud forecast from a weather app with a model trained on the forecast's historical accuracy for a specific location.

This model, called TabPFN v2, reduces the error in the forecast by 28-41% over a year's worth of data from six different cities. The tool provides a verdict for tonight, an hour-by-hour comparison of the app's prediction and the actual chance of clear skies, and a moon phase indicator. Additionally, the author built a local version of the tool that can be trained on personal location data and generate a short plan based on the results.

The project includes a live page for six cities and a sky-photo log that uses Gemma 4 to rate the cloud cover in user-provided photos. The code is available on GitHub, and a demo video showcases the tool's functionality.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

S1MB Número Uno: un juez de cero tokens que lideró el ranking de motores de decisión

Resumen En el ranking de motor de decisión de la System One Mosaic Benchmark (S1MB), nuestro modelo Darwin-27B-ZTC-v2 es número uno entre 102 modelos, con una puntuación Borda de 89.58 y un promedio…

  • Darwin-27B-ZTC-v2 named top decision-making model in S1MB
  • Zero-token generation process sets it apart from competitors
  • Achieved Borda score of 89.58 and average task score of 66.46

One Self-Improving Model, Eleven Number-One Titles: What That Takes

TL;DR One self-improving model family now holds eleven public number-one benchmark records at the same time, across math, science, law, structured output, and decisions.

  • S1MB model breaks records in eleven categories simultaneously
  • Recursive self-improving loop pushes multiple benchmarks to top
  • Zero-token decision method enhances accurate and efficient decision-making

S1MB 1위: 0토큰으로 판정해 디시전 엔진 리더보드를 제패하다

요약 System One Mosaic Benchmark(S1MB) 디시전 엔진 리더보드에서 우리 모델 Darwin-27B-ZTC-v2가 102개 모델 중 1위입니다. Borda 점수 89.58, 태스크 평균 66.46. 핵심은 판정 방식입니다. 한 번의 순전파로, 생성 토큰 0개로 결정합니다. S1MB가 측정하는 것 S1MB는 타입드 결정 품질을 봅니다.

  • Darwin-27B-ZTC-v2 named top decision-making engine on S1MB leaderboard
  • Zero-Token Confidence method uses single forward pass, zero token generation
  • Model achieves 89.58 Borda score, 66.46 task average in S1MB benchmark

Engram Corruption: What Happens When a Skill Container Doesn't Own Its Payload

The setup In a modular AI framework like LivinGrimoire, behavior comes from small, swappable units called skills. A Brain holds them in lobes, and a skill's input() runs on every think cycle.

  • Skills are managed by higher-level AH skills, which handle skill management.
  • Engram snapshot process inadvertently includes payload skills, causing duplication issues.

I built Chloe: an open-source TypeScript framework for AI agents you control

I built Chloe for developers who want to build AI agents for businesses, especially small businesses. The idea is simple: use code for predictable tasks, and use AI where you need it.

  • I developed Chloe, an open-source TypeScript framework for AI agents.
  • Framework uses code for predictable tasks and AI for necessary operations.
  • Chloe provides developers control over agents, maintaining ownership of code.

More from Sunday 11 October →