Adding Low-Fidelity Data to ML Training Sets Could Yield Better Process Models
Low-fidelity trend information should be incorporated into the data sets used to train machine learning-based biopharmaceutical process models, according to new analysis, which suggests the approach could also help identify the most manufacturable drug candidates. The post Adding Low-Fidelity Data to ML Training Sets Could Yield Better Process Models appeared first on GEN - Genetic Engineering…
Researchers suggest that integrating low-fidelity data with high-fidelity experimentally-derived information can enhance the precision of machine learning-based process models in biopharmaceutical development. While machine learning models have the capability to predict the effects of parameter changes pre-emptively, providing the required high-fidelity data for training is costly and scarce.
Lead author Mohammad Golzarijalal, a research fellow at the University of Melbourne’s digital bioprocess hub, proposes incorporating general trends from low-fidelity data into training sets to teach the model the overarching process pattern. The model would then utilize a smaller quantity of high-fidelity data to refine these trends, better aligning the surrogate model predictions with the available ground truth.
Overfitting often occurs when models are trained solely on limited high-fidelity data, resulting in poor performance in unobserved conditions. By learning trends from lower-cost data, a multi-fidelity model can achieve greater predictive accuracy and explore a broader process space without necessitating the same number of expensive experiments.
Such improved models could yield more accurate predictions for cell growth, viability, metabolite concentrations, product titers, and guide various optimization aspects like media, feeding schedules, seeding density, and operating conditions. One example mentioned is a recent study in the researchers' group that merged 20,000 simulated CHO fed-batch data points with data from 65 experimental bioreactor runs.
The multi-fidelity Gaussian-process models demonstrated more accurate predictions of final monoclonal-antibody titer compared to a model trained solely on experimental data. While drug companies currently utilize low-fidelity data to some extent, such as drawing on previous experiments to define parameter ranges, design-of-experiments studies, and process optimization, there is an opportunity to leverage these data more systematically.
Multi-fidelity algorithms offer a structured approach to determine the optimal amount of information to transfer from historical or simulated data to current problems, making past experience more useful for prediction and decision-making.
Written by urgent.news from GEN Biotechnology's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.