Urgent.News

What's breaking now, across thousands of outlets.

AI

Reconnected Claude Code Over SSE — and Found the Distance Scores Aren't as Stable as I Thought

Loose end from Entry 14: switching mcp_search_server.py from stdio to SSE broke the Entry 08 Claude Code connection, which was registered expecting a spawned process, not a network server. Fixing it turned out to be the easy part: claude mcp remove today-i-ran-notes claude mcp add --transport sse today-i-ran-notes http://localhost:8090/sse claude mcp list ✔ Connected , server still running from…

Claude Code's distance scores for retrieving relevant information were found to be less stable than initially thought. When switching from a standard process connection to a network server (SSE) in the code, the connection to the Claude Code entity was disrupted. However, after some adjustments, the connection was successfully restored.

Testing the connection involved asking Claude Code to search for a specific question related to pod status using the 'oc' command. The test was successful, and the tool provided a coherent answer identifying the incorrect command from a previous entry. However, the focus shifted to examining the stability of the distance scores, which proved to be less consistent than expected.

Comparing the distance scores for slightly different phrasings of the same question yielded a significant variance. For instance, the distance between two related results was 54 points, while the gap between the best and worst results within a single query was only 18 points. This indicates that the variance from paraphrasing the question was three times larger than the variance between a highly relevant result and a slightly less relevant one in the same result set.

This finding suggests that the distance metric is not effectively calibrated to determine relevance. The distance score seems to be more influenced by word choice rather than the actual relevance of the information. Additionally, the same query text consistently produced identical distance scores, regardless of the client or transport used. This consistency undermines the assumption that the distance score is a reliable measure of relevance across different queries.

Furthermore, it was discovered that Claude Code's accurate quote of an incorrect command from a previous entry was not generated by the MCP tool. Instead, the tool provided truncated results, and Claude Code used its own file system access to retrieve the complete information. This means that the initial clean-looking answer was not solely a result of the RAG (Retrieval-Augmented Generation) pipeline, but also included Claude Code's direct access to the relevant files on disk.

Lastly, it's worth noting that the collection size in this particular case was relatively small, as every query returned the same five results, just in a different order. The index was not yet populated with a large amount of data, which affects the reliability of the distance rankings. With a smaller dataset, the rankings may not accurately reflect the true relevance of the information.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Friday 2 October →