lisaR: An LLM-Inferred Semantic Annotation of biological categories for gene set enrichment analysis
Gene set enrichment analysis (GSEA) turns differential expression results into lists of enriched gene sets. These lists are often long and redundant and span several gene-set collections, which makes their biological interpretation difficult. We present lisaR, an R package that reads previously computed differential expression or differential abundance results, runs GSEA and organises the output…
Gene set enrichment analysis (GSEA) converts differential expression findings into lists of enriched gene sets. These lists are typically extensive and overlapping, encompassing multiple gene-set collections. This poses challenges for biologists to interpret the biological meaning. To address this issue, a novel R package called lisaR has been developed.
lisaR reads in previously computed differential expression or differential abundance data, performs GSEA, and organizes the output using LISA dictionaries. These dictionaries utilize LLM-Inferred Semantic Annotation to categorize 13,511 gene sets from various sources like Gene Ontology branches, curated pathway collections, and more. The dictionaries are organized into 158 biological categories, grouped into 42 supercategories, and further divided into four dictionaries.
Four open-weight language models evaluate the semantic alignment of each gene set with the categories. A consensus rule then selects the most appropriate assignments. The dictionaries are included with the package and do not rely on data input; hence, analyses conducted at different times with the same dictionaries use consistent categories. Importantly, no language model is queried during the actual analysis.
For each biological category, lisaR summarizes the direction of enrichment in its significant gene sets and performs closed testing using Hommel's robust procedure. This procedure determines whether the category exhibits enrichment evidence and provides a simultaneous lower bound on the number of enriched gene sets within the category.
The results of lisaR are presented in a user-friendly report that can be accessed through a web browser. This report links each category to its corresponding gene sets and genes, with all figures saved along with their associated data and R code. To demonstrate the utility of lisaR, two case studies have been provided. The first involves RNA-seq data from melanoma biopsies before and during nivolumab treatment, while the second features paired tumor and adjacent-tissue proteomics from clear cell renal cell carcinoma.
lisaR is freely available at its GitHub repository (https://github.com/DBM-OlmedaLab/lisaR) and archived at a DOI (https://doi.org/10.5281/zenodo.23044930). Comprehensive documentation is also provided at https://olmedalab.org/lisaR/.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.