PerturbTrace: Evaluating Feedback Use by AI Co-Scientist Agents in Perturbation Discovery
Recent advances in AI co-scientists have brought LLM agents into closed-loop experimental design. However, whether these agents use feedback from earlier rounds to revise subsequent experimental decisions remains unclear. We address this question with PerturbTrace, which evaluates each round-to-round transition through Feedback-to-State, State-to-Action, and Action-to-Outcome. These stages assess…
The research paper "PerturbTrace: Evaluating Feedback Use by AI Co-Scientist Agents in Perturbation Discovery" investigates whether AI co-scientist agents leverage feedback from previous rounds to modify their experimental decisions. To accomplish this, the authors developed PerturbTrace, a tool that assesses the transitions between feedback, state, action, and outcome in each round.
The primary objective is to determine if the AI agents' reasoning and perturbation-selection strategies incorporate feedback, and if the subsequent set of experiments is influenced by the stated strategy, ultimately leading to a higher number of hits compared to random sampling.
The study tested four LLM agents on 17 screen-derived tasks and compared them against random selection, active learning, and LLM-guided Bayesian optimization baselines. The results revealed that each AI agent surpassed the strongest non-agent method on at least 15 of the 17 tasks. However, when conducting controlled evaluations across six tasks, there was no consistent advantage from incorporating true feedback over random or no feedback.
Out of 576 transitions involving true or random feedback, only 43 (7.5%) successfully progressed through the complete Feedback-State-Action-Outcome sequence, including 25 instances under random feedback.
These findings indicate that achieving high final recall does not necessarily imply effective feedback utilization. Furthermore, the research emphasizes the necessity to assess closed-loop scientific agents based on both their discovery performance and the extent to which feedback influences their subsequent decisions.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.