Back to Blog
protein engineering

Why Protein Stability Engineering Matters for Biotech

By Ravit Netzer

Why Protein Stability Engineering Matters for Biotech

Protein instability is not a project risk that shows up dramatically, with clear signals and obvious decision points. It accumulates quietly in background attrition numbers. A campaign to develop a thermostable cellulase for a textile processing application starts with 80 candidate variants, screens 40 for expression, characterizes 20 for activity, and ends with 3 that meet the stability threshold for process evaluation. The team repeats this cycle twice more across different variant families before finding a lead. The biology worked, the chemistry worked, the team executed well. But the economics of three full campaigns to arrive at a single processable lead are very different from the economics of one well-targeted campaign.

Stability is the bottleneck that nobody talks about at the start of a project. This post is about why that happens and what it costs, specifically in the enzyme and early therapeutic protein context where we have the most direct experience.

Why instability is systematically underestimated at project start

When a team picks a target enzyme for engineering, stability is usually not the first design criterion. Function is first: can this enzyme catalyze the reaction we need? Then selectivity and specificity, then expression level, then cost of production. Stability gets framed as a property you will fix in later optimization rounds, after you have a molecule that does what you want biochemically.

The problem is that fixing stability in later rounds is much more expensive than building it in from the start, for two reasons. First, stability engineering and activity engineering are not fully decoupled. Mutations that improve thermostability can reduce activity, and vice versa: the structural rigidity that raises Tm often reduces the backbone flexibility that enzymes need for catalytic turnover. Attempts to improve stability after the activity profile has been locked in have to navigate a constrained search space where improving one property risks degrading the other.

Second, low-stability starting points have a compounding effect on the experimental workflow. An enzyme that denatures rapidly under process conditions requires larger quantities of purified protein per experiment, because activity declines over the assay timescale. Characterization data from unstable proteins is noisier, with higher experimental variance per measurement. Teams doing rational engineering on an unstable parent spend a significant fraction of their effort just managing the experimental artifacts of working with an unstable system.

The attrition pattern in industrial enzyme campaigns

In campaigns we have worked with or reviewed through pilot engagements, a consistent pattern appears: the candidates that fail to advance are disproportionately eliminated by one of three stability-related failure modes rather than by activity failure. The three modes are aggregation under process conditions (the protein forms inactive aggregates when held at elevated temperature for the duration of a process cycle), insufficient half-life at the target operating temperature, and loss of activity under the pH and ionic strength conditions required by the process even at moderate temperatures.

This is not unique to the partners we have worked with. The published literature on industrial enzyme development describes the same attrition pattern: activity leads that fail stability qualification are consistently the dominant source of campaign extension. The discrepancy between the percentage of engineering effort allocated to activity versus stability versus the percentage of attrition caused by each is a genuine mismatch between where teams invest resources and where projects actually fail.

We are not saying that activity engineering is less important than stability engineering: they are both necessary, and neither alone is sufficient. The point is that stability-driven attrition is large, predictable in its structural basis, and more amenable to early computational prediction than most activity failure modes. That asymmetry is the opportunity.

What stability engineering actually involves

Stability engineering in the context of biotech product development typically means one or more of the following: raising the thermal denaturation midpoint (Tm) by 3-10 degrees Celsius above the process operating temperature, extending the functional half-life at operating conditions from hours to days, reducing aggregation propensity under stress conditions, and improving resistance to degradation by the proteases present in the process environment.

Each of these outcomes is achievable through targeted mutagenesis of specific structural features. Tm improvements typically come from changes at buried core positions that improve hydrophobic packing, from the introduction of stabilizing salt bridges at surface positions, and from helix-capping mutations that reduce the entropic cost of the folded state. These mechanisms are well-characterized and predictable in many protein families with adequate structural and evolutionary information.

Aggregation reduction is a different problem: it requires identifying and modifying the surface patches that expose hydrophobic residues prone to intermolecular contacts. This is also computationally tractable, though the prediction accuracy for aggregation is generally lower than for Tm because the relevant process depends on both sequence and concentration, and the interplay with formulation conditions adds a variable that sequence-based models do not capture.

Where computational triage changes the economics

The traditional stability engineering workflow for an enzyme campaign is approximately the following: design a variant panel covering candidate positions based on structural analysis, synthesize 40-80 variants, express and purify in parallel, characterize all expressed variants by DSF or thermal shift assay, and identify hits. The synthesis and characterization phase typically takes 6-10 weeks for a panel of this size, including the cloning and purification steps.

Computational triage changes the input to that synthesis step. Instead of synthesizing 40-80 variants to characterize by DSF and find the hits, you synthesize 8-15 variants that the computational screen has ranked as high-probability hits, characterize all of them, and achieve a similar or better hit rate in a panel one-fifth the size. The economics shift: synthesis cost is reduced, characterization time is reduced, and the iteration cycle shortens from 8 weeks to 2-3 weeks.

There is a ceiling on this benefit. Computational triage is only as good as its prediction accuracy, and prediction accuracy varies substantially by protein family, structure quality, and alignment depth. A 70% hit rate on a 15-variant shortlist is very different from a 70% hit rate on a 60-variant panel in terms of its economic value, but both require the wet-lab confirmation step to be the actual decision point. Computational tools that are used to replace experimental characterization rather than to prioritize it will fail; tools that are used to sharpen the starting list for a characterization campaign that is still required will not.

A practical example from the textile enzyme space

Consider a cutinase development program aimed at biodegradable polyester processing. The target operating temperature is 55 degrees Celsius with a process duration of 4-6 hours. The wild-type cutinase from a mesophilic organism has a Tm around 52 degrees and a half-life at 55 degrees of approximately 45 minutes, making it effectively non-functional under process conditions. The team needs to raise Tm by 5-8 degrees to achieve the half-life target.

A structural scan of a cutinase from this family shows several candidate positions where the evolutionary statistics strongly favor substitutions associated with thermophilic orthologs: three buried positions with suboptimal packing interactions relative to thermophilic homologs, and two surface positions with exposed hydrophobic patches that thermophilic variants consistently reduce. Computational prediction across these positions identifies 4 high-confidence stabilizing substitutions.

Synthesizing and characterizing those 4 variants plus combinatorial combinations (4 singles, 6 doubles, 4 triples, 1 quadruple: 15 variants total) in one round addresses the engineering goal directly. If 2 of the 4 single-position hits validate experimentally with the predicted Tm improvements, the double combination of the best two, and confirmation that cutinase activity is retained, constitutes a project decision point. That is 15 variants and one characterization round. The alternative, unguided scan of all 20 candidate positions with 2 substitutions per position: 40 variants plus combinations, 2-3 characterization rounds.

What this means for where to invest engineering effort early

The practical implication is that stability prediction should be part of parent selection, not just variant optimization. Before committing to a target enzyme as the starting point for an engineering campaign, a computational stability scan of the wild-type and its closest characterized homologs gives you information about how much headroom exists to improve stability, whether the positions likely to carry stability improvements are in regions that also constrain activity, and how many rounds of engineering you are likely to need to reach the process threshold.

This kind of predictive feasibility assessment before committing resources to a multi-round campaign is the highest-leverage use of computational stability prediction. It does not eliminate the experimental work, and it cannot guarantee outcomes. But it changes the decision quality at the front of a project, which is where the most important investment decisions are made.