Back to Blog
computational biology

Computational Stability Screening Versus Experimental Screening

By Ravit Netzer

Computational Stability Screening Versus Experimental Screening

Protein engineers have always done some form of computational pre-screening, even before modern machine learning tools made it systematic. You check the literature for mutagenesis data on homologs, you inspect the structure for packing holes, you look at conservation across the family to see which positions are tolerant of substitution. The question is not whether computation belongs in the workflow. It is which decisions computation can own and which decisions still require measurement.

This post is our attempt to draw that line honestly, based on what we have learned building and testing Scala's stability model. We have a direct interest in making computational screening sound as useful as possible, so we will try to be specific about where it genuinely earns the time savings and where wet-lab screening is still the only reliable answer.

Where computational pre-screening saves the most time

The highest-value application of computational stability screening is variant triage. You have a large candidate space and need to decide which fraction of it to synthesize. The computational screen narrows that fraction based on predicted fitness, so the experimental screen operates on a smaller, more enriched set. The value is multiplicative: if you reduce your synthesis panel from 200 variants to 30 and your hit rate on the computational shortlist is substantially higher than random selection, you have saved both synthesis cost and lab time without losing the actual high performers.

From our pilot program data, this triage function works well when three conditions hold: the target protein has adequate alignment depth (at least several hundred high-quality homologs), the mutations are single-point or small combinatorial sets at surface or semi-surface positions, and the goal is enriching for thermostability improvement rather than predicting exact Tm values. Under these conditions, computational pre-screening is genuinely decision-useful. The shortlist contains the real hits at substantially higher density than a random sample of the same size.

A second strong application is identifying positions that are likely to be intolerant of mutation at all. If a scan consistently predicts negative stability effects for all substitutions at a given position, that is a strong signal that the position is structurally constrained. Taking that signal seriously and leaving those positions alone saves the wasted synthesis and expression labor you would have spent confirming the same thing experimentally.

Where experimental screening remains non-negotiable

Thermostability is one fitness dimension, not all of them. A variant that improves predicted Tm by 5 degrees can still reduce kcat by 40% if the stabilizing substitution constrains active-site flexibility or alters a substrate contact. Enzymatic activity is not captured by the stability model. Full characterization of selected candidates, including kinetic assay under relevant conditions, remains an experimental step that computational screening does not replace.

Solubility and expression titer are also outside the stability prediction scope. A variant that is predicted to be more thermostable may express worse in your E. coli or Pichia system if the substitution affects co-translational folding in ways that the model does not account for. We have seen this in pilot data: a computationally top-ranked variant that showed excellent DSF Tm in the purified form expressed at significantly lower titer than the wild type. The stability was real; the expression problem was not predictable from the sequence-based model.

For any variant that is going beyond the screening stage into scale-up or formulation, experimental characterization of stability, activity, and expression is always necessary. This is not a limitation of Scala's model specifically. It is a fundamental constraint on what thermostability prediction can know from sequence information alone.

The middle ground: where results depend on the specific protein

Between clear wins and clear losses for computational screening sits a category of cases where the value depends strongly on the protein. Proteins with limited homolog representation are the most common source of disappointed expectations. A team working on a cyanobacterial enzyme with 180 high-quality homologs in UniRef90 will get meaningfully less reliable predictions than a team working on a common industrial cellulase family with 15,000 homologs. The confidence tier output communicates this, but it is worth understanding why the reliability varies so much.

Membrane proteins are a structurally distinct category where stability prediction is harder across all current methods, not just ours. The thermodynamic parameters of membrane protein stability are different from soluble proteins, the relevant assays are different, and the evolutionary signals are harder to extract from sequence alignments that are often shorter and more constrained by structural function. We do not recommend using Scala for membrane protein stability prediction in its current form.

Antibody and antibody-fragment engineering is an area where we have had some pilot success but where the engineering goals often blend stability with aggregation propensity and expression in CHO or HEK systems. Pure thermostability prediction is useful as one input, but it is not sufficient as the primary decision criterion for antibody engineering campaigns.

Cost and time comparison: a realistic breakdown

To put some concrete numbers on the comparison: a standard DSF screen of 96 variants, including plate-based expression, clarification, and measurement, typically takes two to four days of bench work for a skilled researcher, not counting synthesis and cloning time. A computational scan of the same 96 positions (all single-point substitutions across those positions) runs in minutes to hours depending on alignment computation. The synthesis and cloning upstream of the experimental screen is the bigger time cost.

If the computational scan can reduce a planned synthesis panel from 96 variants to 24 without losing the actual high performers, the time saving at the synthesis and cloning stage is the dominant benefit. Reducing from 96 to 24 oligos ordered is a meaningful synthesis cost reduction, but the real compression is in the expression, purification, and characterization work. Running DSF on 24 samples instead of 96 is a 4-fold reduction in that bottleneck. For teams that are capacity-constrained at the assay stage, this is where computational pre-screening creates the most direct value.

A decision framework for when to pre-screen computationally

We suggest applying the following criteria before deciding whether to invest in computational pre-screening for a given campaign. First, is the variant space large enough that you need triage? If you are planning to synthesize and test 8 variants based on structural inspection, the computational scan adds less marginal value. If you are looking at 50 or more positions or combinatorial combinations, pre-screening pays off. Second, does your target protein have adequate alignment depth? Check UniRef90 homolog count before running. Third, is thermostability the primary engineering goal, or is it one of several simultaneous objectives? The more it is the primary objective, the more directly useful the computational output is. Fourth, are the mutations within the scope of what the model handles well: single-point substitutions and small double-mutant combinations at non-core positions?

If the answer to most of those questions is yes, computational pre-screening is a good use of the time and cost. If the target protein is a membrane protein, has sparse homologs, or if the engineering campaign is primarily about activity optimization with stability as a secondary constraint, manage expectations accordingly.

The framing we try to maintain internally is: computational screening is a filter layer, not a decision-maker. It compresses the experimental work by enriching the synthesis panel. It does not remove the experimental work.