{
  "id": 12870546,
  "title": "Making LLM ratings auditable: quotes, limitations, and \"can't judge\" instead of zero",
  "url": "https://urgent.news/2026/10/08/making-llm-ratings-auditable-quotes-limitations-and-cant-judge",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-08T12:46:13.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/stiamora/making-llm-ratings-auditable-quotes-limitations-and-cant-judge-instead-of-zero-27in"
  },
  "original_language": "en",
  "account": "I'm creating a side project called WeChat Source. It's a directory aimed at helping Chinese readers discover WeChat Official Accounts, which are essentially newsletters within the WeChat platform. The challenge lies in showcasing LLM-generated ratings without them appearing as confident and misleading. Here are the rules I've established for displaying these ratings.\n\nFirstly, the evidence window is critical. Each account is rated on its 20 most recent articles containing readable text, using no more than the first 6,000 characters of each. If an account has fewer than 20 usable articles, it doesn't receive an overall score. This approach prevents the display of a confident-sounding number without substantial evidence.\n\nSecondly, every rating includes quotes and a limitation section. The five dimensions of evaluation are depth of argument, source transparency, information gain, clarity, and caution in stating opinions. Each dimension receives a 1-5 score, accompanied by a brief explanation, a limitations paragraph highlighting the account's shortcomings in that dimension, and specific quotes from articles as evidence. The page also clarifies the scoring scale: 3 is acceptable, 4 is good, and 5 requires strong evidence.\n\nThirdly, quote verification is crucial. Occasionally, LLMs mistakenly quote non-existent text. These errors are identified by comparing the quoted text against the stored article content. Any incorrect quotes are removed. If this results in insufficient evidence for a dimension, the rating displays as \"can't judge,\" accompanied by a note indicating that the quote verification failed and that a re-analysis or human review is required. Crucially, \"can't judge\" is not treated as a zero score. This approach avoids unfairly penalizing accounts for the model's mistakes.\n\nFourthly, provenance is displayed on each profile. It includes the number of articles used to generate the rating, the model name, the version of the rubric, and the analysis date. Additionally, a line specifies that the score represents the platform's AI analysis, not an official evaluation. This disclaimer is particularly important as older profiles will still be identifiable even if the rubric changes.\n\nLastly, to prevent similarity from being mistaken for quality, a \"keep exploring\" list appears at the bottom of each account page. This list is generated by rule-based matching of shared categories and tags. The page explicitly states this, including the similarity percentage and the shared categories. It's essential to communicate that these recommendations are not quality endorsements but rather a representation of shared content.",
  "summary": "I'm building a side project, WeChat Source . It's a directory that helps Chinese readers find WeChat Official Accounts (the newsletter-style publications inside WeChat). The site is in Chinese, but the design problem behind it isn't: how do you show LLM-generated ratings without them turning into confident-sounding noise? Here are the rules I ended up with. All of them are visible on every…",
  "key_points": [
    "Each account rated on 20 most recent articles with 6,000-character limit",
    "Ratings include quotes, limitations, and clarification of scoring scale",
    "\"Can't judge\" indicates quote verification failure, not zero score"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}