Build your own decision model
Decision models are systems that infer and respond with calibrated probabilities or all possible answers. These models can be found in everyday language models, which may require multiple passes to generate a valid response. Jev is an example of such a decision model that assumes there are fixed options to choose from, allowing for quick selection through a single pass.
By constraining the possible outputs to a specific set, such as A, B, C, D, or E, the model can only emit those tokens. Selecting the highest probability output yields the answer.
However, this does not guarantee the output's correctness. Treating output token probabilities as confidence scores is common, but without additional training, those scores may reflect the model's confidence in the next token rather than the true probability of the correct answer. To test the model's accuracy, run it against public datasets. In this case, the 1.7B model performed reasonably well on a random sample holdout of CommonsenseQA. Finetuning the model led to slightly better performance.
When tested against an ambiguous problem, the model exhibited overconfidence, with its confidence scores not matching its accuracy. The model tended to be extremely overconfident in the 0.9 - 1.0 bin but only correct 70% of the time. When it made predictions with 0.8 - 0.9 confidence, it was only accurate ~40% of the time. This indicates that the model is generally overconfident in its predictions.
To calibrate the model and make the output scores reflect its accuracy, temperature scaling can be used. By adjusting the temperature value, the output probability distribution curve can be flattened and scaled to approximate accuracy. In this case, a temperature value of 3.797280788421631 was found to achieve better calibration. Scripts are available on GitHub to guide users through building a dataset, evaluating, finetuning, and calibrating their own model. It is encouraged to try these methods on other bigger models.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.