AI hackathon: how to test a solution before the final pitch
A team shows a successful model response. A judge changes the input text and gets a different result: a date disappears, an invented city appears or the request hangs. To investigate before the presentation, the organizer needs an evaluation protocol: which examples to test, what to compare the answers against and what to save after each run. Prepare it before development starts. Teams will know…
To effectively evaluate an AI model before the final pitch, an evaluation protocol must be established. This protocol should specify which examples to test, what the results should be compared against, and what to save after each run. The protocol was built for a single AI prototype, designed to extract city, event format, and registration deadline from fictional announcements.
The expected output is JSON with fields city, format, and deadline, with missing values marked as null. The format can be online, offline, or hybrid, and dates should be in YYYY-MM-DD format. Teams are provided with 30 artificial announcements, 18 to use during development and 12 held out for final evaluation. These examples include Russian and Kazakh texts, missing fields, and ambiguous dates.
Before the evaluation, teams and judges need to agree on city name variants, rules for incomplete dates, and handling ambiguous wording. The held-out set should be kept separate and frozen until solutions are finalized. The protocol also includes building a simple baseline using a city dictionary, explicit event format indicators, and recognizing fully specified dates.
After building the baseline, the AI solution should be tested on identical inputs with the same comparison rules. The results should be saved, including the run identifier, commit, held-out dataset version, model name and API version, prompt version, response schema, preprocessing rules, runtime and dependency versions, and generation settings.
A log with one record per input per run should be created, detailing the run ID, example ID, input text, expected fields, raw response, processed response, error type, run status, and reviewer's note. This log should be used to inspect errors example by example.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.