Urgent.News

What's breaking now, across thousands of outlets.

Tech

Code Review Retrieval Explained: Simple Semantic and Keyword Search with Portable Reranking

Short answer: combine lexical and embedding candidates, fuse ranks rather than raw scores, rerank only a small merged set, and refuse to produce a code-review finding unless the final evidence still points to an accessible, current policy passage. This costs more latency than a single search, but it protects the exact identifiers that semantic retrieval tends to blur while keeping the retrieval…

Combine lexical and embedding candidates, then merge their ranks instead of using raw scores. Only rerank a small subset of the merged results and avoid generating any code-review findings unless the final evidence still points to a reachable, current policy section. This approach sacrifices some latency for increased reliability.

The key takeaway is that retrieval success does not equal review success. In a B2B SaaS review pipeline, missing a policy could allow risky changes, while duplicated findings could post the same output multiple times. Every stage should have a unique identifier, a limited retry policy, and evidence that persists through retries.

The example uses a chatbot that searches internal Node.js engineering documentation, evaluates a proposed code change, and produces structured findings. The retrieval service is implemented in Go, though the client language is irrelevant since the same HTTP contract can be called from any language, and the index or reranker can be updated without modifying the review logic.

To combine semantic search, keyword search, embeddings, and reranking, begin with two separate candidate generators. Keyword search should target exact strings like "AbortSignal", "package-lock.json", rule IDs, error codes, and configuration keys. Embedding search should include paraphrases, such as matching "stop work after the caller disconnects" with a policy on cancellation propagation.

Neither branch should be the final decision-maker. Each branch returns document and chunk identities, revision, access scope, rank, and a brief excerpt. Do not merge the raw relevance scores. Instead, merge the rankings using reciprocal rank fusion (RRF), deduplicate based on stable chunk IDs, and rerank the top candidates against both the proposed change and the review question.

Start with 40 candidates from each branch, a merged pool of no more than 60, and 12 inputs for the reranker. These are configuration values, not universal best practices. You may need to adjust them with an offline evaluation set containing exact-token cases, paraphrases, outdated revisions, prohibited documents, and changes that should not produce any finding.

The process is deliberate: verify tenant and repository authorization before conducting any search. Retrieve lexical and semantic candidates simultaneously. Fuse their ranks and eliminate duplicate stable chunk IDs. Rerank a limited pool using the code diff and review question. Validate the revision, authorization, and evidence before generating the final output.

This is the complete retrieval pathway. The initial design that returns a list of strings works in a demo but fails in a production runbook because an excerpt cannot answer critical operational questions: Which document revision created this finding? Was the caller authorized to access it? Were multiple chunks derived from the same policy paragraph?

Can a retry produce the same result? The output must contain all this information. For code review, treat a finding as a derived record with a deterministic identity. Calculate a hash using the repository, commit SHA, policy revision, rule ID, and normalized code location. If a retry occurs, the system can regenerate the record, but the sink should perform an upsert on that key, ensuring that the same key always maps to the same record.

Do not use a newly generated request ID as the deduplication key; it identifies a request, not the actual work. The same principle helps catch missed findings. Record candidate counts for both branches, the fused and reranked counts, reasons for rejection, index revision, policy revision, and end-to-end duration. An empty lexical branch is different from an authorized query with no lexical match.

Similarly, a generator returning no finding is different from a retrieval stage that loses its candidates. Keep these states separate. The explanation is extensive because the root cause of the failure spans multiple boundaries. Imagine a change that replaces a Node.js call accepting an AbortSignal with one that no longer accepts it.

Semantic search finds a general document about resource cleanup, while keyword search finds the exact API policy. After merging, reranking prefers the exact rule in the context of the diff, and the evidence gate ensures that the rule revision is current and readable. The system then emits a structured finding with the rule ID and source chunk.

If the retrieval fails or the exact policy is unavailable, the system should abstain instead of attributing the general cleanup passage to a specific cancellation policy. This example demonstrates the importance of maintaining explicit states and avoiding the creation of false confidence. A portable retrieval core written in Go emphasizes types and test fixtures over specific provider implementations.

Keep provider-specific scores hidden behind two ranked interfaces, maintain stable source metadata, and ensure deterministic fusion. The provided Go package, "retrieval", includes a Query struct to encapsulate search parameters, candidate types for lexical and search results, a Searcher interface for performing searches, and a Reranker interface for reranking candidates.

The Fuse function merges multiple candidate lists while limiting the total number. This approach provides a solid foundation for a retrieval system, ensuring that the evidence contract is maintained throughout the process.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Idempotency: The Secret to Preventing Double Payments and Network Glitches

Idempotency is a property of an operation where executing it multiple times has the exact same effect as executing it once.

  • Idempotency ensures repeated operations yield same result as single execution.
  • Analogy of elevator button illustrates idempotent actions in software.
  • Unique identifiers prevent duplicate charges and data corruption.

More from Tuesday 18 August →