Urgent.News

What's breaking now, across thousands of outlets.

AI

The cheapest way to stop your AI product from regressing

A startup changes the model behind its AI feature. The new model is faster, cheaper and performs better on public benchmarks. The engineering team runs its tests, deploys the update and waits for the improvement. Instead, support tickets begin to arrive. The assistant is less accurate on short questions. It misunderstands customers who mix languages. […] The post The cheapest way to stop your AI…

The cheapest way to stop your AI product from regressing

Generative AI systems can perform well in conventional software testing but may exhibit deteriorating behavior in real-world applications. This issue is exacerbated by the open-ended nature of their outputs, which can vary significantly even for the same input. A common mistake is relying on informal testing with a few examples, which may not reveal the system's true limitations.

To mitigate these risks, a small, carefully curated collection of real user inputs known as a golden dataset is crucial. This dataset should contain a representative sample of user interactions, including ambiguous requests, unsupported languages, and missing information. By running these examples whenever the model or system is updated, potential regressions can be identified and addressed before they affect users.

The golden dataset should not replace other monitoring methods like production monitoring or user feedback, but rather serve as an additional layer of quality assurance. It should be constructed using real production examples, synthetic examples can be used initially, but should be gradually replaced with real interactions. The goal is to capture a recognisable sample of how users actually engage with the product, not to create a perfect benchmark.

When evaluating the AI's performance, focus on the outcome or desired behavior rather than requiring an exact answer. This approach accommodates the variability in human language and ensures the system meets its functional requirements. The dataset should be kept small enough for easy inspection and versioned alongside the product to track changes and improvements.

Written by urgent.news from e27's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at e27.co →

More in AI

Open-PR: một AI agent review PR nói chuyện như đồng nghiệp, không như một con bot

AI coding giúp viết code nhanh hơn, nhưng số PR mở ra cũng nhiều hơn trong khi số người review vẫn vậy. Kết quả: review — chứ không phải viết code — mới là nút thắt cổ chai mới.

  • Open-PR is open-source AI agent based on popular CLI agents
  • Reviews PRs to maintain project's language and style, not temporary comments
  • Accesses project docs, memory for each repo to avoid repeating previous feedback

OpenAI Cracks a Million-Dollar Math Problem — and the Credit Fight Starts Immediately

OpenAI Cracks a Million-Dollar Math Problem — and the Credit Fight Starts Immediately Today's biggest AI story isn't really about the breakthrough itself — it's about who gets to claim it.

  • OpenAI claims internal model solved Navier-Stokes problem, a Millennium Prize challenge.
  • Mathematicians dispute OpenAI's claim, say they had similar approach ready earlier.
  • OpenAI denies data access, acknowledges potential influence of general usage patterns.

More from Thursday 10 September →