Skip to main content

Write a comment

PREreview of Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization

Published
DOI
10.5281/zenodo.22711725
License
CC BY 4.0

This manuscript introduces MLEE, a machine-learning-assisted extension of htFuncLib that uses experimentally measured activity data to redirect stability-enriched libraries toward a substrate-specific function. The workflow is benchmarked across three distinct fitness landscapes and experimentally demonstrated for MthUPO, in which the second-round library substantially increases the fraction of variants that exceed WT activity. Overall, the protocol appears robust and data-efficient, with repeated benchmark simulations and the observation that using four versus eleven screening plates yielded essentially the same MthUPO library-design decision. 

Minor comments:

  1. Could the authors provide a table listing the top-ranked predicted variants, including their full mutational sequences, predicted activities, measured activities where available, and indicate the rank of 11C5? This would clarify the basis for selecting 11C5 and provide a more direct assessment of how well the model prioritizes individual high-activity sequences.

    1. The rationale for selecting MthUPO_11C5 for detailed characterization could be outlined in the text. 11C5 was identified experimentally during the initial htFuncLib screening and subsequently characterized as the lead variant; however, the manuscript does not show how 11C5 compares with the highest-ranked sequences predicted by the activity model. 

    2. Although the authors report that the MLEE-enriched library contains 9 of the top 10 and 74 of the top 100 predicted variants, the identities and predicted/measured activities of these variants are not provided. 

    3. While the text outlines the prevalence of double, triple, quadruple, and quintuple mutants that exceeded WT activity (lines 325-330), it is unclear if these higher activity mutants all derive their high activity from different epistatic pairs or if they derive from the same one or two base mutations 

  2. Can the authors provide more detailed information on the library complexity and coverage for both rounds?

    1. The manuscript reports that 11 screening plates yielded 440 unique variants after sequencing, which were subsequently used to train the activity prediction model. It also demonstrates that training on the first four plates reproduces essentially the same MLEE-enriched library. However, it is difficult to assess the amount of information required for model training because the manuscript does not report how many unique sequence-activity pairs were obtained from the first four plates after removing duplicate variants. Including the number of unique variants used for training and their distribution in SI in both the 4-plate and 11-plate analyses would help readers better understand the data efficiency of the proposed workflow.

    2. The methods section mentions verification of the library (line 526) but does not report the results or provide an analysis of library size and coverage prior to assaying. 

  3. The manuscript adopts a continuous fitness-dependent sample weighting scheme to emphasize the comparatively sparse high-activity variants during regression model training. However, we found it difficult to determine whether the improved library enrichment is primarily due to the weighting scheme, the AAindex representation, or the position-specific library design strategy. It could be useful to include an ablation comparing weighted regression or the strategy introduced in Ref. 41.

  4. The authors repeatedly emphasized in the manuscript that this regression model is used only to derive position-specific amino acid rankings, not to predict measured variant activities. It would be helpful to discuss the rationale for choosing a regression objective over ranking-based learning objectives.

  5. The supervised regression model is trained on multi-mutant variants and therefore has the potential to capture epistatic sequence–activity relationships; however, the subsequent library-design strategy reduces these sequence-level predictions to WT-background single-residue preferences. It would be helpful to further evaluate whether this projection preserves the benefit of learning from combinatorial sequence data:

    1. Residue preferences are currently inferred by predicting each substitution in the WT background. Because the effect of a substitution can depend strongly on its genetic context, the authors could compare this approach with scoring each residue across multiple genetic backgrounds. For example, by averaging its predicted activity over variants containing that residue. This would test whether incorporating background-dependent effects provides a more robust residue ranking.

    2. The manuscript motivates constructing a residue-level combinatorial library rather than directly selecting the highest-scoring predicted variants. Since the regression model predicts activities for the complete htFuncLib sequence space, it would be informative to compare the proposed residue-ranking strategy against a simpler baseline in which the top 144 predicted variants are selected directly for second-round screening. While GGAssembly results in pooled libraries, a fragment design to account for more of the epistatic signal would not significantly improve the experimental design, as it would more directly test paired mutational effects. Comparison of these experimentally matched designs would help justify the additional projection from sequence-level predictions to residue-level rankings.

Competing interests

The authors declare that they have no competing interests.

Use of Artificial Intelligence (AI)

The authors declare that they did not use generative AI to come up with new ideas for their review.

You can write a comment on this PREreview of Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now