Back to Blog
epistasis

Epistatic Interactions in Protein Mutation Panels

By Tamar Levy

Epistatic Interactions in Protein Mutation Panels

A common failure mode in protein engineering campaigns goes like this: you identify two mutations that each improve thermostability individually, you combine them in a double mutant expecting additive or better-than-additive improvement, and the double mutant is worse than either single. Sometimes it is worse than wild type. You have invested synthesis and characterization resources on a variant that the single-point predictions would have confidently ranked near the top, and it fails because of an interaction that single-point analysis cannot see.

This is a description of negative epistasis, and it is not rare. Estimates from deep mutational scanning literature suggest that 20 to 40% of double mutants combining individually beneficial substitutions show some degree of epistatic antagonism. The range is large because the frequency depends strongly on the structural and functional relationships between the positions being combined. Understanding why this happens, and what computational approaches can do to help, requires getting into the mechanism.

The mechanistic basis of epistasis in stability engineering

Epistasis in protein stability arises because residues are not independent. The free energy contribution of a substitution at position A depends on the context set by every other position in the protein, including position B. When B is wild type, substitution at A may relieve strain in a loop by changing local backbone geometry. When B is mutated simultaneously, the loop geometry is already different, and the same change at A may now introduce a new clash or disrupt a recently formed hydrogen bond.

The clearest cases of negative epistasis involve physically proximal positions where packing rearrangements propagate between residues. If position A and position B are within about 6 to 8 angstroms of each other, a substitution at A changes the packing environment for B's sidechain, and the beneficial substitution at B was optimized against the original A environment. The combination is epistatic almost by design.

But epistasis also operates at longer range through electrostatic networks, allosteric coupling, and changes to the overall stability margin. A stabilizing mutation that increases thermal resistance by removing conformational strain may reduce the protein's ability to accommodate a second mutation that also requires conformational flexibility. The structural intuition for long-range epistasis is harder to develop and harder to reason about from a crystal structure alone.

Why additive models underperform for multi-site prediction

Most rapid combinatorial screening workflows assume that the effects of multiple mutations are approximately additive: the predicted stability of a double mutant is the sum of the two individual delta-Tm predictions. This assumption is computationally convenient and works reasonably well when the positions are structurally distant and the individual effects are small. It breaks down systematically when positions are structurally coupled, when individual effects are large, or when the background sequence is already at a stability extreme.

In our early pilot data, we tracked the accuracy of additive prediction for double mutants where we had experimental data for both singles and the double. The additive model predicted the correct sign (stabilizing versus destabilizing) for about 78% of double mutants. That sounds acceptable until you consider that the 22% misclassifications are disproportionately concentrated in the most interesting cases: large-effect combinations and proximal-position pairs. If you are only combining your top two single-point hits, and those happen to be proximal positions with large individual effects, the additive model is exactly the place where it is most likely to mislead you.

How Scala approaches pairwise and higher-order interactions

Our approach to epistasis prediction has two components. The first is based on co-evolutionary analysis. Positions that co-vary in multiple sequence alignments across a protein family are likely to be functionally or structurally coupled. When two positions show strong co-evolutionary signal, their combined substitution effects are more likely to be non-additive. We use direct coupling analysis (DCA) derived features to flag position pairs with high co-evolutionary coupling as candidates for non-additive prediction treatment.

The second component is a learned interaction term in the prediction model. Rather than predicting double-mutant stability as a simple sum of single-mutant predictions, the model learns pairwise interaction terms from the training data. Where experimental data for double mutants exists in our training set for a given protein family, those interaction terms are directly calibrated. For protein families without double-mutant data in the training set, the interaction terms are transferred from the co-evolutionary coupling features.

The limitation is important: our interaction term predictions are more reliable for well-characterized families where double-mutant experimental data exists in the training set. For sparse families, we are relying on the DCA-derived signal, which is noisier. We communicate this through the confidence tier output: position pairs flagged as having strong co-evolutionary coupling but low direct calibration data get a higher uncertainty annotation on their double-mutant predictions.

Practical implications for how you build a multi-site campaign

The most direct implication is that you should treat physically proximal position pairs differently from distant-position pairs when planning combinatorial synthesis. For positions within roughly 8 angstroms of each other in the structure, individual single-point hits should not be automatically combined. Scrutinize the co-evolutionary coupling signal for those pairs, and if there is any indication of strong coupling, prioritize synthesizing the double mutant as an explicit test rather than assuming additivity.

For distant positions, the additive assumption is substantially more reliable. A double mutant combining a strongly predicted stabilizing substitution at position 45 with an equally strong prediction at position 200, where the two positions are far apart in the structure and show no co-evolutionary coupling, will behave additively in the large majority of cases.

A second practical implication involves the order of operations in iterative engineering campaigns. If you are planning to stack multiple substitutions over several rounds of engineering, the epistatic context changes with each round. A substitution that was neutral or slightly positive on the original wild-type background may become strongly positive or strongly negative on a background that already carries two earlier-round mutations. Our combinatorial scan accounts for this by computing predictions against the specified background sequence, not necessarily the original wild type. Getting the background right is the user's responsibility: you need to specify which background sequence you are scanning against when setting up the run.

The limits of epistasis prediction at scale

Higher-order interactions (triple and beyond) are currently outside reliable computational reach for most protein families. The number of possible triple combinations grows too fast, and the experimental data needed to calibrate higher-order interaction terms at scale does not exist for most families. Deep mutational scanning experiments that cover double mutants exhaustively for a protein of typical size are already large and expensive; triple-mutant coverage is impractical for any but the smallest target regions.

Our practical recommendation is to treat higher-order combinations with caution regardless of what the summed single-point predictions suggest. When combining three or more individually beneficial substitutions, always include single and double intermediates in the synthesis panel so you can measure the interaction landscape at each step rather than going straight to the fully combined variant. That experimental data, fed back into subsequent rounds, is the best available calibration for the specific epistatic landscape of your protein.