PhageLysData: an evidence-aware and AI-ready dataset of phage lytic enzymes and depolymerases
Bacteriophage lytic enzymes and depolymerases are relevant to phage biology, antimicrobial development, and protein engineering, but their sequence and annotation data remain dispersed across general databases, specialized resources, genome-centred collections, and prediction-oriented datasets. We present PhageLysData, an evidence-aware and AI-ready resource constructed through reproducible…
PhageLysData is a new dataset designed to bring together information on phage lytic enzymes and depolymerases in a comprehensive and organized manner. This resource is particularly useful for those in the fields of phage biology, antimicrobial development, and protein engineering, as it consolidates data that has been scattered across various databases and collections.
PhageLysData is built through a systematic integration of information from seven primary resources, ensuring that the data is reliable and traceable. The dataset consists of 807,366 source observations that have been combined into 759,105 unique exact-sequence entities. This includes a Core set of 11,867 entities, which are supported by evidence, a Prediction Extension of 745,092 sequences that are derived from predictions, and 2,146 Context entities that provide provenance and reference information.
The Core entities in PhageLysData are enriched with a wealth of information, including harmonized biological annotations, physicochemical properties, functional annotations derived from InterProScan, mapped structural assets from PDB and AlphaFold DB, and reusable numerical representations. For a substantial subset of 11,259 eligible Core sequences, the dataset also provides embeddings from 11 different protein language models, as well as one-hot encoding, all of which are represented using a common contract.
To demonstrate the capabilities of PhageLysData, the release includes examples that show how it can be used for exploring latent spaces, performing unsupervised clustering, conducting supervised classification, and retrieving evidence-aware candidates. Importantly, these examples do not rely on a universal predictive benchmark, allowing for a more flexible and adaptable approach to working with the data.
Overall, PhageLysData serves as a solid foundation for a variety of applications involving phage lytic enzymes and depolymerases. It supports tasks such as protein retrieval, comparative analysis, creating task-specific datasets, and developing machine-learning applications. The dataset is designed to be versioned and computationally accessible, ensuring that it remains a valuable resource for researchers in the field.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.