A 6-Case Single API Key Acceptance Harness for Compatible SaaS Chat
A single API key saves deployment work, but it does not make model behavior portable. For a SaaS that scores candidates against a job rubric, the useful unit of portability is a versioned scoring contract backed by six acceptance cases. My choice is one small adapter boundary, validation on every response, and promotion only after a provider-model pair clears that harness. Short answer: compare…
A single API key simplifies deployment, but it lacks portability. For scoring candidates against a job rubric, the unit of portability is a versioned scoring contract, backed by six acceptance cases. The goal is to compare integrations based on application behavior, not credentials. Both OpenAI, Claude, and Gemini can be integrated through a common chat-shaped request, yet they remain distinct execution targets.
A compatible chat-shaped request represents transport compatibility. Stable candidate decisions are a product requirement.
Maintaining a small adapter boundary, validation on every response, and promotion only after a provider-model pair passes the harness ensures consistent scoring. Outsourcing secret storage and schema validation is recommended, as they are undifferentiated tasks. The rubric and decision rules should remain in application code.
Candidate scoring raises complex questions about rubric completion, score evidence, and handling of missing evidence. These questions cannot be resolved solely by accepting requests. Each model must be recorded as a tested target, and behavioral equivalence should not be inferred from provider names.
The scoring process should start with scoreCandidate(), not a generic chat() method. Batch work deserves a separate lifecycle. The OpenAI Batch API guide outlines uploading input, creating batches, checking status, and retrieving output. However, a portable application should model deferred jobs separately, distinct from blocking chat calls.
The scoring contract should carry domain facts, not provider-specific vocabulary. It should handle unknown evidence as null and zero for failed criteria. Combining these creates false precision. The contract should include Criterion, ScoreRequest, CriterionResult, ScoreResult, and ScoringTarget interfaces.
Returning unknown is intentional. Provider output becomes application data only after local validation. A schema library is advisable, but checks must be domain requirements. Fail closed and send missing evidence to manual review, rather than converting it to zero.
The final weighted score should be computed locally, with the application owning the weights. The function weightedScore() determines the weighted score based on request and result. Six cases before production traffic ensure the scoring contract works for all cases. These cases cover scenarios such as strong evidence for every criterion, no evidence for one criterion, contradicting evidence, resume text containing scoring instructions, long but valid work history, and near-identical candidates around the threshold. Running every fixture against each provider-model pair before promotion is necessary.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.