Turning Domain Expertise into Multi-Dimensional Evaluation of Biomedical AI with Karenina
Language models and agents are increasingly used in biomedicine, but current benchmarks reward correct answers even when the underlying reasoning is flawed. Here we introduce Karenina, an open-source framework that turns expert knowledge into multi-dimensional evaluations of questions, conversations and autonomous agents. Illustrated in Question-Answer pairs, multi-turn conversations and…
We haven't written up this one. bioRxiv has the full story — the link below goes straight to it.