-
PDF
- Split View
-
Views
-
Cite
Cite
Samir Char, Nathaniel Corley, Sarah Alamdari, Kevin K Yang, Ava P Amini, ProtNote: a multimodal method for protein-function annotation, Bioinformatics, 2025;, btaf170, https://doi.org/10.1093/bioinformatics/btaf170
- Share Icon Share
Abstract
Understanding the protein sequence-function relationship is essential for advancing protein biology and engineering. However, less than 1% of known protein sequences have human-verified functions. While deep learning methods have demonstrated promise for protein function prediction, current models are limited to predicting only those functions on which they were trained.
Here, we introduce ProtNote, a multimodal deep learning model that leverages free-form text to enable both supervised and zero-shot protein function prediction. ProtNote not only maintains near state-of-the-art performance for annotations in its training set, but also generalizes to unseen and novel functions in zero-shot test settings. ProtNote demonstrates superior performance in prediction of novel GO annotations and EC numbers compared to baseline models by capturing nuanced sequence-function relationships that unlock a range of biological use cases inaccessible to prior models. We envision that ProtNote will enhance protein function discovery by enabling scientists to use free text inputs without restriction to predefined labels – a necessary capability for navigating the dynamic landscape of protein biology.
The code is available on GitHub: https://github.com/microsoft/protnote; model weights, datasets, and evaluation metrics are provided via Zenodo: https://zenodo.org/records/13897920.
Supplementary Information is available at Bioinformatics online.
Author notes
Current affiliation for Nathaniel Corley Institute for Protein Design, University of Washington, Seattle, WA, USA 98195