收藏切换
VenusMutHub: A systematic evaluation of protein mutation effect predictors on small-scale experimental data
收藏切换
PDF
Liang Zhanga, d, Hua Pangb, c, Chenghao Zhange, Song Lia, Yang Tana, c, d, Fan Jianga, d, Mingchen Lia, c, d, Yuanxi Yua, d, Ziyi Zhoua, d, Banghao Wua, e, Bingxin Zhoua, Hao Liuf, Pan Tana, c, d, *, Liang Honga, d, e, *
Acta Pharmaceutica Sinica B | 2025, 15(5) : 2454 - 2467
Less
收藏切换
Acta Pharmaceutica Sinica B | 2025, 15(5): 2454-2467
TOOLS
VenusMutHub: A systematic evaluation of protein mutation effect predictors on small-scale experimental data
Full
Liang Zhanga, d, Hua Pangb, c, Chenghao Zhange, Song Lia, Yang Tana, c, d, Fan Jianga, d, Mingchen Lia, c, d, Yuanxi Yua, d, Ziyi Zhoua, d, Banghao Wua, e, Bingxin Zhoua, Hao Liuf, Pan Tana, c, d, *, Liang Honga, d, e, *
Affiliations
  • aSchool of Physics and Astronomy & Institute of Natural Sciences, Shanghai Jiao Tong University, Shanghai National Centre for Applied Mathematics (SJTU Center), MOE-LSC, Shanghai 200240, China
  • biHuman Institute and School of Life Sciences and Technology, Shanghaitech University, Shanghai 201210, China
  • cZhangjiang Institute for Advanced Study, Shanghai Jiao Tong University, Shanghai 201203, China
  • dShanghai Artificial Intelligence Laboratory, Shanghai 200232, China
  • eSchool of Life Sciences and Biotechnology, Shanghai Jiao Tong University, Shanghai 200240, China
  • fShanghai Matwings Technology Co., Ltd., 200240 Shanghai, China
About Author:

E-mail addresses: (Pan Tan),

These authors made equal contributions to this work.

Author contributions

Liang Zhang: Writing – original draft, Software, Visualization, Methodology, Data curation. Hua Pang: Data curation. Chenghao Zhang: Data curation, Visualization. Song Li: Resources. Yang Tan: Resources. Fan Jiang: Resources. Yuanxi Yu: Resources. Mingchen Li: Methodology. Banghao Wu: Methodology. Bingxin Zhou: Methodology. Ziyi Zhou: Methodology. Liang Hong: Review & editing, Project administration, Funding acquisition, Supervision. Pan Tan: Writing – review & editing, Project administration, Funding acquisition, Supervision.

doi: 10.1016/j.apsb.2025.03.028
Outline
收藏切换

In protein engineering, while computational models are increasingly used to predict mutation effects, their evaluations primarily rely on high-throughput deep mutational scanning (DMS) experiments that use surrogate readouts, which may not adequately capture the complex biochemical properties of interest. Many proteins and their functions cannot be assessed through high-throughput methods due to technical limitations or the nature of the desired properties, and this is particularly true for the real industrial application scenario. Therefore, the desired testing datasets, will be small-size (∼10–100) experimental data for each protein, and involve as many proteins as possible and as many properties as possible, which is, however, lacking. Here, we present VenusMutHub, a comprehensive benchmark study using 905 small-scale experimental datasets curated from published literature and public databases, spanning 527 proteins across diverse functional properties including stability, activity, binding affinity, and selectivity. These datasets feature direct biochemical measurements rather than surrogate readouts, providing a more rigorous assessment of model performance in predicting mutations that affect specific molecular functions. We evaluate 23 computational models across various methodological paradigms, such as sequence-based, structure-informed and evolutionary approaches. This benchmark provides practical guidance for selecting appropriate prediction methods in protein engineering applications where accurate prediction of specific functional properties is crucial.

Protein engineering  /  Mutation effect prediction  /  Benchmark  /  Small-scale experimental data  /  Stability  /  Activity  /  Binding affinity  /  Selectivity
Liang Zhang, Hua Pang, Chenghao Zhang, Song Li, Yang Tan, Fan Jiang, Mingchen Li, Yuanxi Yu, Ziyi Zhou, Banghao Wu, Bingxin Zhou, Hao Liu, Pan Tan, Liang Hong. VenusMutHub: A systematic evaluation of protein mutation effect predictors on small-scale experimental data[J]. Acta Pharmaceutica Sinica B, 2025 , 15 (5) : 2454 -2467 . DOI: 10.1016/j.apsb.2025.03.028
Protein engineering has emerged as a powerful approach for developing proteins with enhanced or novel functions, playing crucial roles in various applications ranging from industrial biocatalysis to therapeutic protein development1-3. The significance of mutation effect prediction is particularly profound in pharmaceutical science, where understanding protein mutations is essential for drug development and precision medicine4. In biopharmaceutical development, accurate prediction of mutation effects can guide the optimization of protein therapeutics for improved stability, manufacturability, and clinical efficacy5,6. Moreover, in precision medicine, understanding the impact of missense mutations in drug targets is crucial for predicting patient-specific drug responses and developing targeted therapeutic strategies7. However, traditional wet-lab experimental methods are often limited by their throughput and resource requirements when screening the vast mutational sequence space, making it impractical to experimentally test all possible mutations. This limitation has highlighted the critical role of computational approaches in protein engineering. A central challenge in this field is the accurate prediction of mutation effects, which holds the potential to significantly accelerate protein design workflows by prioritizing promising candidates for experimental validation8,9.
Recent years have witnessed remarkable progress in zero-shot computational methods for predicting protein mutation effects. These approaches span from physics-based methods10 to modern machine learning models11, offering diverse strategies for mutation effect prediction. Deep learning models have shown significant advancements in predicting the effects of protein mutations, leveraging their capabilities in understanding complex sequence-function relationships12. However, most of these models have been primarily evaluated on large-scale mutation datasets, which derived from deep mutational scanning experiments13. While such evaluation provides valuable insights into overall model performance, they have a fundamental limitation: to achieve high-throughput screening, DMS experiments must rely on approximate assays and surrogate readouts (e.g., fluorescence-based assays or growth-based selections) that may not adequately capture or reflect the actual biochemical properties of interest14,15. This limitation becomes particularly critical when the target properties involve complex molecular functions or require specific physiological contexts that can hardly be approximated by high-throughput screening methods.
In real-world protein engineering scenarios, particularly in directed evolution applications, researchers often employ an iterative approach where initial experimental data from a small number of carefully selected mutations guide subsequent rounds of engineering. These mutations are typically chosen from functionally important regions based on structural knowledge or mechanistic understanding. Many proteins of interest cannot be subjected to high-throughput screening methods due to technical limitations, cost constraints, or the nature of the desired function16. Consequently, available experimental data often consists of small-scale datasets, typically ranging from dozens to several hundred mutations. The first-round experimental data not only validates initial predictions but also provides crucial information for selecting the most suitable prediction model for the next round of engineering. This creates a significant gap between the evaluation conditions of prediction models and their practical application contexts, where accurate prediction of mutations in key regions is essential for efficient iterative optimization.
Furthermore, protein engineers need to predict various functional properties including stability, catalytic activity, binding affinity and selectivity. Different prediction models may exhibit varying performance across these properties, and their effectiveness might depend on the size and quality of available training data. Given these challenges, the evaluation of models on small-scale datasets is particularly crucial because: (1) it reflects the real constraints in industrial and academic protein engineering projects16, (2) it enables assessment using direct biochemical measurements of the target properties rather than surrogate readouts14, and (3) it provides insights into model selection for the next round of engineering based on initial experimental data. Despite the abundance of prediction methods, there lacks a systematic evaluation of their performance on small-scale datasets that better represent real-world protein engineering scenarios through rigorous experimental characterization.
Motivated by these practical challenges, we present VenusMutHub, a comprehensive benchmark study evaluating diverse mutation effect prediction models. We systematically collected and curated 905 small-scale experimental mutation datasets from published literature and public databases, spanning 527 proteins across various functional properties, including stability, activity, binding affinity, and selectivity. Our evaluation encompasses diverse prediction approaches including protein language models, structure-aware models, and alignment-based methods, assessing their performance across different dataset sizes and functional properties. This benchmark aims to provide practical insights for protein engineers by objectively assessing model performance under realistic economic and technical constraints.
To address the challenges outlined in the introduction and provide a comprehensive evaluation of protein mutation effect prediction models, we established a systematic benchmark workflow as illustrated in Fig. 2. Our approach consists of four main steps: dataset building from published literature and databases, data processing, model prediction using 23 different computational approaches, and performance evaluation using various metrics. The following sections detail each component of this workflow.
Data collection We compiled our benchmark datasets from published literature and public databases, with a focus on protein engineering applications. For stability measurements (ΔΔG, ΔTm), we integrated several established datasets including FireProtDB and ThermoMutDB17,18, along with other curated stability datasets (S571, S4346, S8754, M126119, S66920, S264821). For binding affinity data, we curated two types of datasets: protein–protein interaction (PPI) data from PPB-Affinity22, focusing on single-point mutations occurring on a single chain, and drug–target interaction (DTI) data from BindingDB23, including Kd, Ki, and IC50 measurements for mutations of the same wild-type protein. We also curated catalytic activity (kcat, kcat/km, km, specific activity and vmax) and selectivity (e.e., d.e.) data from published protein engineering studies. To ensure data quality, we only included datasets with clear documentation of experimental methods and quantitative measurements.
Dataset characteristics As shown in Fig. 1A, our benchmark encompasses 905 distinct datasets from 527 unique proteins, with stability measurements constituting the majority (59.7%, 540 datasets), followed by activity (19.3%, 175 datasets), binding (15.8%, 143 datasets), and selectivity (5.2%, 47 datasets). The protein sequence length distribution (Fig. 1B) spans from small peptides to large proteins (0–1400 residues). The distribution of dataset sizes (Fig. 1C) reflects typical protein engineering scenarios, where most experiments involve a limited number of mutations: 578 datasets are very small (5–20 mutations), 197 are small (21–50 mutations), 77 are medium-sized (51–100 mutations), and 53 datasets contain more than 100 mutations. This size distribution aligns well with real-world protein engineering practices, where comprehensive deep mutation scanning is often impractical.
Data processing To ensure data quality and consistency, we implemented a comprehensive preprocessing pipeline. We applied strict quality control criteria by removing mutations where the annotated wild-type residue did not match the corresponding position in the reference sequence (e.g., a mutation annotated as M1H was removed if the first residue in the wild-type sequence was not methionine). Additionally, we required each dataset to contain at least 5 mutations with corresponding experimental measurements to ensure sufficient data points for analysis. To address potential redundancy, we developed a deduplication process using MD5 hash values based on mutation and score information. For binding affinity measurements (Kd, Ki, IC50) where lower experimental values indicate higher affinity, we implemented a crucial value transformation step. These values were converted to -log10 scales prior to model training and evaluation, ensuring alignment with the model's scoring convention where higher values represent superior mutant performance. The processed datasets were standardized to a consistent format containing mutant identifiers, mutation sequences, and experimental measurements, facilitating uniform model evaluation across all datasets.
For structural information required by various prediction models, we utilized ColabFold24 to generate both multiple sequence alignments (MSAs) and protein structure predictions. The MSAs generated by ColabFold in a3m format were converted to a2m format using the reformat.pl script from HH-suite325 to accommodate the input requirements of various prediction models. This standardized pipeline ensured consistent structural and evolutionary information across all proteins in our benchmark.
We evaluated a comprehensive set of 23 protein mutation effect prediction models, which can be categorized based on their input features.
These models only require the target protein sequence as input:
ESM ESM-1b11, ESM-1v26 and ESM-227 are protein language models with transformer encoder architecture similar to BERT, trained with masked language modeling objectives on UniRef50 or UniRef90. They predict fitness using the masked-marginal approach, which provides optimal performance on substitutions.
CARP CARP28 is a protein language model trained with MLM objective on UniRef50. The architecture leverages convolutions instead of self-attention, leading to computational speedups while maintaining high downstream task performance.
RITA RITA29 is an autoregressive language model akin to GPT230, trained on UniRef100 with model sizes ranging from 85M to 1.2B parameters.
ProGen2 ProGen231 is an autoregressive protein language model trained on a mixture of UniRef90 and BFD3032, with models ranging from 151M to 6.4B parameters. It employs a standard transformer decoder architecture.
ProtGPT2 ProtGPT233 is a 738M-parameter autoregressive transformer trained on Uniref50, capable of generating de novo protein sequences that explore novel regions of the protein space while maintaining natural protein-like properties.
UniRep UniRep34 trains a Long Short-Term Memory (LSTM) model on UniRef50 sequences.
These models leverage evolutionary information through different approaches:
GEMME GEMME35 is a Global Epistatic Model that predicts mutational effects by analyzing multiple sequence alignments. It combines evolutionary tree-based conservation analysis with site-wise frequencies to calculate epistatic effects across the entire sequence, providing a statistical approach to mutation effect prediction.
VESPA VESPA36 is a Single Amino acid Variant (SAV) effect predictor that combines embeddings from the protein language model ProtT5 with per-residue conservation predictions.
MSA transformer MSA Transformer37 directly processes multiple sequence alignments using a specialized transformer architecture designed to capture dependencies both within and across sequences in the alignment.
Tranception Tranception38 offers flexibility in utilizing MSA information through its autoregressive architecture, with variants available for both MSA-based and sequence-only prediction.
PoET PoET39 takes an innovative approach by modeling protein families using unaligned homologous sequences with an autoregressive architecture.
VenusREM VenusREM40 integrates sequence, structure, and evolutionary information through a unified transformer architecture. It processes protein structures via a structure tokenization module, while incorporating evolutionary information through an Alignment Tokenization Module, and combines these features using disentangled multi-head attention for mutation effect prediction.
These models incorporate structural information in different ways:
MULAN MULAN41 builds upon ESM-2 architecture by adding a Structure Adapter module to incorporate backbone angle information, enabling effective integration of structural features while maintaining the power of pretrained language models.
ProSST ProSST12 is a structure-aware protein language model that consists of two key components: a GVP42-based structure quantization module that encodes local residue environments into discrete tokens, and a transformer with sequence-structure disentangled attention that explicitly models the relationship between sequence and structural features. The model is pre-trained on 18.8 million protein structures using masked language modeling.
ProtSSN ProtSSN43 combines semantic and geometric encoding to capture both sequence and structural information. It uses ESM-2 for semantic encoding of protein sequences and Equivariant Graph Neural Networks (EGNN)44 for geometric encoding of protein structures. The model is trained through self-supervised learning by recovering protein information from noisy sequences and structures.
SaProt SaProt45 introduces a structure-aware vocabulary that combines residue and 3D geometric information through Foldseek46-based structure tokenization. This 650M-parameter model converts protein sequences into structure-aware tokens that can be processed by standard language model architectures.
Inverse folding models MIFST47, MIF47, ProteinMPNN48 and ESM-IF149 are inverse folding models specifically designed to model the relationship between protein sequence and structure, generating sequences compatible with given backbone structures.
All models were evaluated using their publicly available implementations and pretrained weights. For models requiring structural information in single-chain prediction tasks, we utilized protein structures generated by ColabFold50. For protein–protein interaction (PPI) predictions, ESM-IF1 and ProteinMPNN were evaluated using experimental complex structures from the PDB database, as these models can process multi-chain inputs. Multiple sequence alignments, when required, were also generated using the ColabFold pipeline to ensure consistency across evaluations.
To comprehensively assess model performance, we employed four key evaluation metrics that capture different aspects of prediction accuracy. The primary metric we used is the Spearman correlation coefficient (ρ) between predicted and experimental values as shown in Eq. (1):
ρ=16i=1ndi2n(n21)
where di is the difference between the ranks of corresponding predicted and experimental values, and n is the number of mutations. Spearman correlation evaluates models' ability to correctly rank mutations regardless of the absolute scale, which is crucial for prioritizing mutations in protein engineering applications. To assess models' ability to identify beneficial mutations across the entire dataset, we calculated the Normalized Discounted Cumulative Gain (NDCG) as shown in Eqs. (2) and (3):
NDCG=DCGIDCG
DCG=i=1p2reli1log2(i+1),IDCG=i=1p2reli1log2(i+1)
where reli represents the relevance score calculated as follows: first, mutations are ranked based on model predictions; the i th relevance score is calculated as its experimental value normalized to [0,1] using min–max normalization as shown in Eq. (4):
reli=vivminvmaxvmin
where vi is the experimental value of mutation i, and vmin and vmax are the minimum and maximum experimental values across all mutations. reli* represents these same relevance scores sorted in descending order (the ideal ranking), and p is the total number of mutations. NDCG is particularly suitable for evaluating mutation prioritization as it measures how well a model's top predictions align with their actual experimental improvements.
For binary classification metrics, we established the classification threshold based on available data characteristics. For datasets containing wild-type fitness measurements, mutations with fitness scores higher than the wild-type were classified as positive samples, while those lower were classified as negative samples. For datasets without wild-type measurements, we used the median fitness score as the threshold. Similarly, predictions were classified as positive if their predicted scores exceeded the corresponding threshold (wild-type prediction or median predicted score). Based on these criteria, we calculated accuracy as shown in Eq. (5):
Accurarcy=TP+TNTP+TN+FP+FN
The F1 score provides a balanced measure between precision and recall, which is particularly important when beneficial mutations are rare as shown in Eqs. (6) and (7):
Precision=TPTP+FP,Recall=TPTP+FN
F1=2·Precision·RecallPrecision+Recall
These complementary metrics provide a comprehensive evaluation: Spearman correlation captures the overall ranking ability, NDCG focuses on identifying top mutations, while accuracy and F1 score assess the models’ capability in distinguishing beneficial mutations from deleterious ones.
We constructed a comprehensive mutation benchmark dataset and systematically evaluated the performance of 23 zero-shot protein mutation fitness prediction models, revealing several key insights about their predictive capabilities (Table 1). VenusREM, which integrates sequence, structure, and MSA information, achieves the highest Spearman correlation and NDCG on both datasets (0.230 and 0.863 for all mutations; 0.285 and 0.878 for single-point mutations). Meanwhile, MIF emerges as another top performer, achieving the highest scores in accuracy and F1 metrics (0.612 and 0.547 for all mutations; 0.627 and 0.562 for single-point mutations), while maintaining competitive performance in ranking metrics (second highest Spearman correlation of 0.229 for all mutations and 0.261 for single-point mutations).
Structure-aware models, particularly inverse folding models (MIF, MIFST, ESM-IF1) and other structure-based approaches (VenusREM, ProSST), consistently rank among top performers in our evaluation. This apparent advantage of structure-aware models, however, should be interpreted with caution as it may be partially attributed to the composition of our dataset, where stability-related mutations constitute a significant portion. We will examine this potential bias in subsequent sections through detailed functional analysis. The strong performance of VenusREM further suggests that the integration of multiple information sources (sequence, structure, and evolutionary information) can provide complementary signals for mutation effect prediction. In particular, we observe distinct performance patterns between the Spearman correlation and NDCG metrics. While VenusREM achieves the highest Spearman correlation (0.285) and NDCG (0.878) for single-point mutations, likely benefiting from its comprehensive information integration, MIF shows exceptional performance in accuracy (0.627) and F1 score (0.562). These variations suggest that different models may excel at capturing different aspects of the mutation landscape.
To ensure the reliability of our analysis, we constructed two datasets: a complete dataset containing both single and multiple mutations (comprising 905 individual datasets), and a refined single-point mutation dataset (comprising 796 individual datasets) by removing all multiple-mutation entries and retaining only those assays with at least 5 single-point mutations. Notably, all models show substantially better performance on the single-point mutation dataset compared to the complete dataset, revealing a universal challenge in capturing epistatic effects. This challenge stems from several fundamental limitations in current modeling approaches. First, the current transformer-based architectures typically predict multiple mutation effects through simple combination of individual mutation impacts, which does not account for non-additive epistatic interactions where one mutation can dramatically alter the structural and functional context for subsequent mutations. Second, the vast combinatorial space of multiple mutations means that most possible combinations are not represented in training data, making it difficult for models to learn generalizable patterns of mutation interactions. These limitations pose significant challenges for protein engineering applications, as multiple simultaneous mutations are often required for achieving desired functional improvements. Therefore, to ensure reliable model assessment, our subsequent analyses focus exclusively on single-point mutations, while acknowledging that improving multiple-mutation predictions remains a critical area for future development. Future advances may come from integrating molecular dynamics simulations to capture long-range epistatic interactions mediated through networks of correlated protein dynamics, or from developing specialized architectures that can explicitly model the cross-correlations between mutation sites Table 1: Overall Model Performance (905 datasets for all mutations and 796 datasets for single-point mutations. The metrics for each dataset are calculated and the values represent the average across all datasets).
Stability predictions, comprising the largest subset of our benchmark (N = 498, 62.5% of total datasets), revealed clear patterns in model performance that align with our understanding of protein stability mechanisms. Given that protein stability is inherently linked to structural characteristics, it is notable that structure aware models, particularly inverse folding models, demonstrated superior predictive capabilities in this category. MIF emerged as the top performer across all evaluation metrics (Table 2), achieving a Spearman correlation of 0.369, NDCG of 0.898, accuracy of 0.620, and F1 score of 0.599. Other structure-aware models also showed strong performance, with ESM-IF1 and MIFST achieving Spearman correlations of 0.343 and 0.336 respectively. Furthermore, VenusREM's strong performance (Spearman correlation 0.355, NDCG 0.897, accuracy 0.613) demonstrates that integrating multiple information sources (structural, evolutionary, and sequence information) can provide complementary signals for stability assessment, achieving the second-best performance across most metrics.
In activity predictions (N = 130, 16.3% of total datasets), we observed distinct performance patterns across different model types. VESPA, which combines conservation predictions (from ProtT5 embeddings), BLOSUM62 substitution scores, and context-aware substitution probabilities, achieved the highest performance in ranking metrics (Table 3), with a Spearman correlation 0.338 and NDCG of 0.833. Models directly utilizing MSA information, such as GEMME (Spearman correlation = 0.277, NDCG = 0.828), Tranception (Spearman correlation = 0.273), and VenusREM (Spearman correlation =0.291, NDCG = 0.831), also demonstrated strong predictive power across different evaluation metrics. The competitive performance of these evolution-informed approaches, including PoET (Spearman correlation = 0.302) which utilizes homologous sequence information, suggests that evolutionary information, whether used directly or indirectly, is particularly valuable for predicting activity-altering mutations. Meanwhile, inverse folding models achieved the highest accuracy scores, with MIFST and MIF reaching accuracies of 0.721 and 0.717 respectively, while Tranception excelled in identifying beneficial mutations as indicated by its top F1 score of 0.559. This diverse performance pattern suggests that different modeling approaches may capture complementary aspects of protein activity mechanisms. The superior accuracy of inverse folding models indicates their strength in structure-aware binary classification, while evolution-informed models excel at capturing the nuanced effects of mutations through ranking metrics. This observation highlights the potential benefit of integrating both structural and evolutionary information for activity prediction, as demonstrated by the balanced performance of models like VenusREM across different metrics (Spearman correlation = 0.291, NDCG = 0.831).
We separated binding affinity predictions into protein–protein interactions (PPI, N = 100, 12.6% of total datasets) and drug–target interactions (DTI, N = 42, 5.3% of total datasets), considering their distinct molecular mechanisms and interaction interfaces (Table 4). This separation revealed interesting patterns in model performance across different binding types. In PPI predictions, most models showed relatively modest performance, suggesting the inherent complexity of protein–protein binding interfaces. Notably, the multichain versions of ESM-IF1 and ProteinMPNN significantly outperformed their single-chain counterparts: ESM-IF1 (multichain) achieved Spearman correlation of 0.193 and NDCG of 0.875 compared to its single-chain version (Spearman correlation = 0.049, NDCG = 0.856), while ProteinMPNN (multichain) showed similar improvements (Spearman correlation 0.130 vs 0.034), demonstrating the importance of incorporating partner protein structural information for accurate PPI predictions. VESPA also achieved competitive performance across all metrics (Spearman correlation = 0.148, NDCG = 0.881, accuracy = 0.572, F1 score = 0.551), indicating that the combination of conservation patterns and substitution probabilities can effectively capture protein–protein binding effects even without explicit partner information.
For DTI predictions, models generally demonstrated better performance, with less pronounced differences between different modeling approaches. The top performers included structure-aware SaProt (Spearman correlation 0.271), MSA-based GEMME (Spearman correlation 0.240), and PoET (NDCG 0.887), which utilizes homologous sequences. All these models achieved competitive results across different metrics, with PoET notably reaching an accuracy of 0.756 and F1 score of 0.623. This balanced performance suggests that drug–protein binding might be more effectively captured by current computational approaches, possibly due to the more localized nature of protein-small molecule interactions.
Selectivity predictions (N = 26, 3.3% of total datasets), consisting of enantioselectivity (ee) and diastereoselectivity (de) measurements, presented the most challenging category in our evaluation. The performance across all models was notably poor, with most models achieving near-zero or even negative Spearman correlations (Table 5). Even the best performing models showed only marginally better results than random predictions: SaProt achieved the highest Spearman correlation of 0.099, while ProtGPT2 and ProSSN followed with correlations of 0.087 and 0.078 respectively. The classification metrics were similarly disappointing, with the highest accuracy of 0.570 (ProteinMPNN) and F1 score of 0.579 (UniRep) being substantially lower than those observed in other prediction tasks.
This universally poor performance can be attributed to several fundamental limitations of current prediction paradigms. First, selectivity measurements inherently involve the relative preference between different substrates (such as different enantiomers or diastereomers), while current models only consider protein information without explicitly modeling substrate-specific interactions. The ee and de values are derived from the relative concentrations of different products, making it particularly challenging for models to predict such differential binding or catalytic preferences without access to small molecule information. Second, the stereochemical nature of selectivity requires understanding of precise three-dimensional arrangements and chemical interactions, which current sequence-based and even structure-based models may not capture adequately.
To address these limitations, future improvements in selectivity prediction may require several key developments: (1) integration of protein-ligand docking simulations to explicitly model substrate binding modes, (2) incorporation of stereochemistry-aware features that can capture the spatial relationships determining enantioselectivity and diastereoselectivity, and (3) development of specialized architectures that can process both protein and substrate information in a manner that accounts for their relative spatial orientations. While the small dataset size (N = 26) poses an additional challenge, the fundamental issue appears to be the mismatch between current modeling approaches and the inherent nature of selectivity prediction tasks.
To investigate how dataset size affects model performance in protein mutation effect prediction, we analyzed 23 representative models using 905 mutation datasets. To ensure balanced sample distribution, we divided the datasets into four quartile-based groups based on mutation counts: 5–8 mutations (234 datasets), 8–13 mutations (236 datasets), 13–28 mutations (215 datasets), and > 28 mutations (220 datasets).
As shown in Fig. 3, we categorized models into three groups based on their underlying principles: structure-aware models, sequence-only models, and evolution-informed models. Structure-aware models generally demonstrated strong performance scaling, with VenusREM and MIF achieving the highest correlations (ρ ≈ 0.36 and ρ ≈ 0.35) on large datasets (>28 mutations). Notably, ESM-IF1 showed comparable performance (ρ ≈ 0.34), suggesting the effectiveness of structure-based approaches. In the sequence-only category, ESM2-650M emerged as the top performer (ρ ≈ 0.29), followed by ESM1v and ESM1b, though some models like UniRep showed limited effectiveness (ρ < 0). Evolution-informed models displayed remarkable consistency, with VenusREM achieving the highest correlation (ρ ≈ 0.36) and other models (GEMME, MSA Transformer, PoET) reaching ρ ≈ 0.30 on large datasets.
These findings suggest that model selection should consider both the available experimental data size and the desired prediction approach. Structure-aware and evolution-informed models generally provide more reliable predictions across different dataset sizes, while sequence-only models show greater performance variation and may require larger datasets (>13 mutations) for optimal performance. The most substantial improvements typically occur between the 5–8 and 8–13 mutation groups, indicating a minimum threshold for reliable predictions.
In addition to average performance metrics, the consistency of model predictions across different datasets is crucial for practical applications. Fig. 4 shows the average standard deviation of Spearman correlation coefficients for different model types across functional categories, providing insights into prediction stability.
Our analysis reveals significant variations in prediction stability across different model types and functional categories. For stability predictions, structure-aware models demonstrate lower variance (σ ≈ 0.35) compared to sequence-only models (σ ≈ 0.39), suggesting that incorporating structural information leads to more reliable stability predictions across diverse proteins. In contrast, for activity predictions, evolution-informed models show more consistent performance (σ ≈ 0.29) compared to both structure-aware (σ ≈ 0.32) and sequence-only models (σ ≈ 0.30). This pattern indicates that evolutionary information provides more stable signals for activity-related mutations across different protein families. The highest variance was observed in selectivity predictions for both evolution-informed and structure-aware models (σ ≈ 0.51), highlighting the particularly challenging nature of this prediction task. Notably, sequence-only models showed surprisingly lower variance in selectivity predictions (σ ≈ 0.25), though this may be influenced by their generally lower performance overall. For binding affinity predictions, structure-aware models exhibited lower variance in protein–protein interactions (σ ≈ 0.30) compared to evolution-informed models (σ ≈ 0.36), indicating that structural information provides more consistent signals for PPI predictions. However, for drug–target interactions, sequence-only models showed competitive consistency (σ ≈ 0.32) compared to structure-aware models (σ ≈ 0.39).
These findings emphasize that model selection should consider not only average performance but also prediction consistency. In practical applications where reliability is paramount, models with lower performance variance may be preferable even if they do not achieve the highest average scores. Furthermore, this analysis highlights the need for uncertainty quantification methods in mutation effect prediction, which would allow practitioners to make more informed decisions by considering both the predicted effect and the confidence level of that prediction.
In this study, we constructed a comprehensive mutation benchmark dataset encompassing 905 distinct datasets across stability, activity, binding, and selectivity properties, and systematically evaluated 23 protein mutation effect prediction models. Our analysis revealed distinct performance patterns: structure-aware models, particularly inverse folding models like MIF and MIFST, demonstrated superior performance in stability predictions, while evolution-informed models like VESPA and GEMME excelled in activity predictions. VenusREM, which integrates sequence, structure, and evolutionary information, showed strong overall performance across multiple metrics. For binding affinity predictions, we observed different patterns between protein–protein interactions and drug–target interactions, with models incorporating partner protein structural information showing advantages in PPI predictions. Notably, all models showed substantially better performance on single-point mutations compared to multiple mutations, and selectivity predictions emerged as the most challenging category.
Our analysis of dataset size impact further revealed that model performance generally improves with larger datasets, with structure-aware and evolution-informed models providing more reliable predictions across different dataset sizes. We identified a critical threshold around 8–13 mutations, where most models show substantial performance improvements, suggesting a minimum data requirement for reliable predictions. Additionally, our performance variance analysis demonstrated that different model types exhibit varying levels of prediction consistency across functional categories, with structure-aware models showing more reliable stability predictions, while evolution-informed models provide more consistent activity predictions. These findings highlight that model selection should consider not only average performance but also prediction consistency and the available experimental data size.
These findings have significant implications for both protein engineering and pharmaceutical science. In biopharmaceutical development, the superior performance of structure-aware models in stability prediction and the distinct patterns in drug–target interaction predictions can guide the selection of appropriate computational tools in different stages of drug discovery pipelines. The identified challenges in predicting protein–protein interactions and the advantages of multichain models highlight important considerations for developing therapeutic antibodies and other protein-based drugs. However, several limitations should be considered: the uneven distribution across different protein properties (stability-related mutations comprising 62.5% while selectivity-related only 3.3%), potential sampling bias in protein families, and the dataset's composition might not fully reflect the diverse challenges in practical applications.
Looking forward, advancing mutation effect prediction will require both architectural innovations and better integration of different information sources. The complementary strengths of different modeling approaches suggest that hybrid or ensemble methods might offer promising directions for future development, particularly for challenging properties like multiple mutations, selectivity, and protein–protein interactions. Furthermore, incorporating uncertainty quantification methods would enhance the practical utility of these models by providing reliability measures alongside predictions. These insights can help both protein engineers and pharmaceutical scientists make more informed decisions in their computational workflows, especially in the early stages of development where experimental validation is costly and time-consuming.
1.
Lutz S, Samantha MI. Protein engineering: past, present, and future. In: Bornscheuer UT, Höhne M, editors. Protein engineering: methods and protocols. New York, NY: Springer; 2018. p. 1—12.
2.
Arnold FH. Directed evolution: bringing new chemistry to life. Angew Chem Int Ed 2018;57:4143—8.
3.
Kim S, Ga S, Bae H, Lee EY. Multidisciplinary approaches for enzyme bio-catalysis in pharmaceuticals: protein engineering, computational biology, and nanoarchitectonics. EES Catal 2024;2:14—48.
4.
Dugger SA, Platt A, Goldstein DB. Drug development in the era of precision medicine. Nat Rev Drug Discov 2018;17:183—96.
5.
Ebrahimi SB, Samanta D. Engineering protein-based therapeutics through structural and chemical design. Nat Commun 2023;14:2411.
6.
Frokjaer S, Otzen DE. Protein drug stability: a formulation challenge. Nat Rev Drug Discov 2005;4:298—306.
7.
Emmerich CH, Gamboa LM, Hofmann MC, Bonin-Andresen M, Arbach O, Schendel P, et al. Improving target assessment in biomedical research: the GOT-IT recommendations. Nat Rev Drug Discov 2021;20:64—81.
8.
Yang KK, Wu Z, Arnold FH. Machine-learning-guided directed evolution for protein engineering. Nat Methods 2019;16:687—94.
9.
Sanavia T, Birolo G, Montanucci L, Turina P, Capriotti E, Fariselli P. Limitations and challenges in protein stability prediction upon genome variations: towards future applications in precision medicine. Comput Struct Biotechnol J 2020;18:1968—79.
10.
Alford RF, Leaver-Fay A, Jeliazkov JR, O’Meara MJ, DiMaio FP, Park H, et al. The Rosetta all-atom energy function for macromolecular modeling and design. J Chem Theor Comput 2017;13:3031—48.
11.
Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci U S A 2021;118:e2016239118.
12.
Li M, Tan Y, Ma X, Zhong B, Yu H, Zhou Z, et al. ProSST: protein language modeling with quantized structure and disentangled attention. NeurIPS 2024;37:35700—26.
13.
Notin P, Kollasch A, Ritter D, Niekerk L, Paul S, Spinner H, et al. Proteingym: large-scale benchmarks for protein fitness prediction and design. NeurIPS 2023;36:64331—79.
14.
Fowler DM, Fields S. Deep mutational scanning: a new style of protein science. Nat Methods 2014;11:801—7.
15.
Wei H, Li X. Deep mutational scanning: a versatile tool in systematically mapping genotypes to phenotypes. Front Genet 2023;14:1087267.
16.
Arnold FH. Innovation by evolution: bringing new chemistry to life (Nobel Lecture). Angew Chem Int Ed 2019;58:14420—6.
17.
Stourac J, Dubrava J, Musil M, Horackova J, Damborsky J, Mazurenko S, et al. FireProtDB: database of manually curated protein stability data. Nucleic Acids Res 2021;49:D319—24.
18.
Xavier JS, Nguyen TB, Karmarkar M, Portelli S, Rezende PM, Velloso JP, et al. ThermoMutDB: a thermodynamic database for missense mutations. Nucleic Acids Res 2021;49:D475—9.
19.
Xu Y, Liu D, Gong H. Improving the prediction of protein stability changes upon mutations by geometric learning and a pre-training strategy. Nat Comput Sci 2024;4:840—50.
20.
Pancotti C, Benevenuta S, Birolo G, Alberini V, Repetto V, Sanavia T, et al. Predicting protein stability changes upon single-point mutation: a thorough comparison of the available tools on a new dataset. Brief Bioinf 2022;23:bbab555.
21.
Dehouck Y, Grosfils A, Folch B, Gilis D, Bogaerts P, Rooman M. Fast and accurate predictions of protein stability changes upon mutations using statistical potentials and neural networks: PoPMuSiC-2.0. Bioinformatics 2009;25:2537—43.
22.
Liu H, Chen P, Zhai X, Huo KG, Zhou S, Han L, et al. PPB-Affinity: protein-protein Binding Affinity dataset for AI-based protein drug discovery. Sci Data 2024;11:1—11.
23.
Gilson MK, Liu T, Baitaluk M, Nicola G, Hwang L, Chong J. BindingDB in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic Acids Res 2016;44:D1045—53.
24.
Mirdita M, Schütze K, Moriwaki Y, Heo L, Ovchinnikov S, Steinegger M. ColabFold: making protein folding accessible to all. Nat Methods 2022;19:679—82.
25.
Steinegger M, Meier M, Mirdita M, Vöhringer H, Haunsberger SJ, Söding J. HH-suite3 for fast remote homology detection and deep protein annotation. BMC Bioinf 2019;20:1—15.
26.
Meier J, Rao R, Verkuil R, Liu J, Sercu T, Rives A. Language models enable zero-shot prediction of the effects of mutations on protein function. Adv Neural Inform Process Syst 2021;34:29287—303.
27.
Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023;379:1123—30.
28.
Yang KK, Fusi N, Lu AX. Convolutions are competitive with transformers for protein sequence pretraining. Cell Syst 2024;15:286—94.
29.
Hesslow D, Zanichelli N, Notin P, Poli I, Marks D. Rita: a study on scaling up generative protein sequence models. arXiv 2022. Available from: https://doi.org/10.48550/arXiv.2205.05789.
30.
Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I, et al. Language models are unsupervised multitask learners. OpenAI blog 2019;1:9.
31.
Nijkamp E, Ruffolo JA, Weinstein EN, Naik N, Madani A. Progen2: exploring the boundaries of protein language models. Cel Syst 2023;14:968—78.
32.
Steinegger M, Söding J. Clustering huge protein sequence sets in linear time. Nat Commun 2018;9:2542.
33.
Ferruz N, Schmidt S, Höcker B. ProtGPT2 is a deep unsupervised language model for protein design. Nat Commun 2022;13:4348.
34.
Alley EC, Khimulya G, Biswas S, AlQuraishi M, Church GM. Unified rational protein engineering with sequence-based deep representation learning. Nat Methods 2019;16:1315—22.
35.
Laine E, Karami Y, Carbone A. GEMME: a simple and fast global epistatic model predicting mutational effects. Mol Biol Evol 2019;36:2604—19.
36.
Marquet C, Heinzinger M, Olenyi T, Dallago C, Erckert K, Bernhofer M, et al. Embeddings from protein language models predict conservation and variant effects. Hum Genet 2022;141:1629—47.
37.
Rao RM, Liu J, Verkuil R, Meier J, Canny J, Abbeel P, et al. MSA transformer. ICML PMLR; 2021. p. 8844—56.
38.
Notin P, Dias M, Frazer J, Marchena-Hurtado J, Gomez AN, Marks D, et al. Tranception: protein fitness prediction with autoregressive transformers and inference-time retrieval. In: International conference on machine learning. PMLR; 2022. p. 16990—7017.
39.
Truong Jr T, Bepler T. Poet: a generative model of protein families as sequences-of-sequences. Adv Neural Inform Process Syst 2023;36:77379—415.
40.
Tan Y, Wang R, Wu B, Hong L, Zhou B. Retrieval-enhanced mutation mastery: agmenting zero-shot prediction of protein language model. arXiv preprint arXiv:241021127 2024. Available from: https://doi.org/10.48550/arXiv.2410.21127.
41.
Frolova D, Pak M, Litvin A, Sharov I, Ivankov D, Oseledets I. MULAN: multimodal protein language model for sequence and structure encoding. bioRxiv 2024. Available from: https://doi.org/10.1101/2024.05.30.596565.
42.
Jing B, Eismann S, Suriana P, Townshend RJ, Dror R. Learning from protein structure with geometric vector perceptrons. arXiv 2020. Available from: https://doi.org/10.48550/arXiv.2009.01411.
43.
Tan Y, Zhou B, Zheng L, Fan G, Hong L. Semantical and topological protein encoding toward enhanced bioactivity and thermostability. Elife 2023:RP98033.
44.
Satorras VG, Hoogeboom E, Welling M. E(n) equivariant graph neural networks. In: International conference on machine learning. PMLR;2021. p. 9323—32.
45.
Su J, Han C, Zhou Y, Shan J, Zhou X, Yuan F. Saprot: potein language modeling with structure-aware vocabulary. bioRxiv 2023. Available from: https://doi.org/10.1101/2023.10.01.560349.
46.
van Kempen M, Kim SS, Tumescheit C, Mirdita M, Gilchrist CL, Söding J, et al. Foldseek: fast and accurate protein structure search. Nat. Biotechnol 2024;42:243—6.
47.
Yang KK, Zanichelli N, Yeh H. Masked inverse folding with sequence transfer for protein representation learning. Protein Eng Des Sel 2023;36:gzad015.
48.
Dauparas J, Anishchenko I, Bennett N, Bai H, Ragotte RJ, Milles LF, et al. Robust deep learning—based protein sequence design using ProteinMPNN. Science 2022;378:49—56.
49.
Hsu C, Verkuil R, Liu J, Lin Z, Hie B, Sercu T, et al. Learning inverse folding from millions of predicted structures. In: International conference on machine learning. PMLR; 2022. p. 8946—70.
50.
Mirdita M, Schütze K, Moriwaki Y, Heo L, Ovchinnikov S, Steinegger MJ. ColabFold: making protein folding accessible to all. Nat Methods 2022;19:679.
Year 2025 volume 15 Issue 5
PDF
10
7
Cite this Article
BibTeX
Article Info
doi: 10.1016/j.apsb.2025.03.028
  • Receive Date:2025-02-13
  • Online Date:2026-09-17
Article Data
Affiliations
History
  • Received:2025-02-13
  • Revised:2025-03-02
  • Accepted:2025-03-05
Affiliations
    aSchool of Physics and Astronomy & Institute of Natural Sciences, Shanghai Jiao Tong University, Shanghai National Centre for Applied Mathematics (SJTU Center), MOE-LSC, Shanghai 200240, China
    biHuman Institute and School of Life Sciences and Technology, Shanghaitech University, Shanghai 201210, China
    cZhangjiang Institute for Advanced Study, Shanghai Jiao Tong University, Shanghai 201203, China
    dShanghai Artificial Intelligence Laboratory, Shanghai 200232, China
    eSchool of Life Sciences and Biotechnology, Shanghai Jiao Tong University, Shanghai 200240, China
    fShanghai Matwings Technology Co., Ltd., 200240 Shanghai, China

Corresponding:

* Corresponding authors.
References
Share
https://castjournals.cast.org.cn/joweb/apsb/EN/10.1016/j.apsb.2025.03.028
Share to
QR

Scan QR to access full text

Cite this article
BibTeX
Citations
表12种不同金属材料的力学参数

Family
属数
Number of
genus
种数
Number of
species
占总种数比例
Percentage of
total species (%)

Genus
种数
Number of
species
占总种数比例
Percentage of total
species (%)
鹅膏菌科Amanitaceae 2 11 5.26 鹅膏菌属 Amanita 10 4.78
小菇科 Mycenaceae 2 12 5.74 丝盖伞属 Inocybe 5 2.39
多孔菌科 Polyporaceae 8 14 6.70 蜡蘑属 Laccaria 5 2.39
红菇科 Russulaceae 3 23 11.00 小皮伞属 Marasmius 6 2.87
小菇属 Mycena 11 5.26
光柄菇属 Pluteus 5 2.39
红菇属 Russula 17 8.13
栓菌属 Trametes 5 2.39
关闭全屏
  • BibTeX
  • EndNote
  • RefWorks
  • TxT