
Table of contents
When a virus mutates to resist the drugs designed to kill it, the consequences can be devastating. HIV drug resistance remains one of the most pressing challenges in infectious disease medicine, and at the heart of understanding it lies a deceptively complex question: how do proteins evolve? The mathematical models researchers use to answer that question, known as substitution models, have long shaped our understanding of molecular evolution. Yet most of these models carry a fundamental blind spot. They treat all sites in a protein as evolving under the same pressures, ignoring the biochemical reality that some positions are under intense functional constraint while others are relatively free to vary. A new study published in Molecular Biology and Evolution proposes a more biologically honest solution, one that could change how researchers model protein evolution in medically critical contexts.
What Are Substitution Models and Why Do They Matter?
Substitution models describe the rates at which one amino acid replaces another over evolutionary time. In phylogenetics, they are the statistical engine behind reconstructing evolutionary trees and inferring ancestral sequences. Without them, we cannot reliably determine how proteins have changed across species or lineages, nor can we accurately predict what ancestral proteins looked like before modern variants emerged.
For decades, the field has relied on empirical substitution models, such as WAG, LG, and JTT, which are derived from large databases of aligned protein sequences. These models estimate average substitution rates across many proteins and have proven remarkably useful. However, they come with a significant limitation: they assume that every site in a protein evolves at the same rate and under the same selective pressures. In reality, a residue buried in the active site of an enzyme faces very different evolutionary constraints than one on the protein surface. Empirical models smooth over this site-to-site variability, which can introduce systematic errors into phylogenetic inference and ancestral reconstruction, particularly for functionally constrained proteins.
The Missing Piece: Enzymatic Activity
Researchers have made progress by developing structurally constrained substitution models, which incorporate information about protein three-dimensional structure. By accounting for factors such as solvent exposure and contact density, these models better capture why certain residues are conserved. They represent a genuine improvement over purely empirical approaches.
Yet even structurally constrained models leave a critical gap. Knowing the shape of a protein does not tell you whether it actually works. Enzymatic activity, the ability of a protein to catalyse a specific biochemical reaction, is not simply a product of structure. It depends on dynamic properties: how flexibly the protein moves, how tightly it binds its substrate, and how efficiently it converts that substrate into product. For enzymes like proteases, which are central to viral replication and are prime drug targets, ignoring enzymatic activity means ignoring the very selection pressures that most strongly shape their evolution. A model that cannot account for whether a protein functions well will systematically misrepresent the evolutionary landscape of functionally critical regions.
A New Model That Combines Structure and Function
The 2024 paper by Ferreiro, Khalil, Sousa, and Arenas introduces a substitution model that goes beyond structure to incorporate enzymatic activity directly. The model quantifies functional constraints through a set of biophysically meaningful descriptors, all derived from molecular dynamics (MD) simulations:
- Binding affinity of the enzyme-substrate complex: how strongly the protease binds its target peptide, a direct proxy for catalytic efficiency.
- Flexibility of structural flaps: the dynamic movement of the flap regions that gate access to the active site, which is essential for substrate entry and product release.
- Hydrogen bonds: the network of interactions stabilising the protein-substrate interface.
- Amino acid backbone radius of gyration: a measure of the compactness of the protein backbone, reflecting overall structural integrity.
- Solvent-accessible surface area: the degree to which residues are exposed to solvent, a classical structural constraint now combined with functional data.
By grounding these descriptors in MD simulations, the model captures the dynamic, time-averaged behaviour of the protein rather than relying on a single static structure. This makes it substantially more biologically realistic than any previous approach, because it reflects the actual physical chemistry that natural selection acts upon.
Testing It on HIV-1 Protease
HIV-1 protease is an ideal test case for this kind of model. It is one of the most thoroughly characterised enzymes in molecular biology, a critical drug target for AIDS treatment, and a protein whose function is absolutely essential for viral replication. Mutations in HIV-1 protease that confer drug resistance are well documented, making it a rich system for studying the interplay between structural constraint, functional requirement, and evolutionary change.
When Ferreiro and colleagues applied their new model to HIV-1 protease sequence data, the results were clear. The activity-incorporating model outperformed both the best-fitting empirical substitution model and a structurally constrained model that does not account for enzymatic activity. The advantage was particularly pronounced in datasets with high molecular identity, where sequences are closely related and subtle site-specific variation is most informative. Crucially, accounting for selection on protein activity improved the fit between modelled evolutionary patterns and real observed sequences in functionally important regions of the protein, precisely the regions where drug resistance mutations tend to arise.
Key Findings
- Empirical substitution models miss important site-specific variation in protein evolution
- Structurally constrained models improve realism but still overlook enzymatic function
- The new model integrates both structural and activity-based constraints for a more complete picture of protein evolution
- Binding affinity and molecular dynamics simulations provide quantifiable functional constraints that can be incorporated into evolutionary models
- The model shows superior phylogenetic likelihood compared to existing empirical and structurally constrained approaches
- Improvements are especially pronounced in datasets with high molecular identity, where site-specific signals are strongest
- The approach is applicable beyond HIV-1 protease to other enzymes where functional constraints can be quantified through MD simulations
What This Means for PhD Researchers
For researchers working at the intersection of computational biology, phylogenetics, and drug resistance, this paper opens a meaningful new avenue. If your work involves modelling the evolution of enzymes, reconstructing ancestral sequences, or understanding how resistance mutations emerge and spread, the framework presented here offers a more realistic foundation than anything previously available. The ability to incorporate MD-derived functional descriptors into substitution models means that the evolutionary pressures most relevant to your biological question can now be explicitly represented in your analysis.
TopPhDs regularly covers cutting-edge computational research that is reshaping how scientists approach complex problems. Whether you are deep in your dissertation or navigating the challenges of a PhD, staying current with methodological advances like this one is part of what separates good research from great research. For bioinformaticians and computational biologists in particular, the integration of molecular dynamics into phylogenetic modelling represents a convergence of disciplines that is likely to become increasingly standard practice.
Paper Reference
Ferreiro D, Khalil R, Sousa SF, Arenas M. Substitution Models of Protein Evolution with Selection on Enzymatic Activity. Mol Biol Evol. 2024;41(2):msae026.
Read the full open-access paper on Oxford Academic.
The message from this research is clear: if you are modelling protein evolution and you are not accounting for enzymatic activity, you are leaving important biological signal on the table. Incorporating activity-based constraints into substitution models is not merely a technical refinement; it is a step toward analyses that are genuinely grounded in the biology of how proteins live, function, and evolve. For researchers working on medically relevant enzymes, that step could make the difference between a model that approximates reality and one that truly illuminates it.
