The standard assumption in computational protein stability work is that you need a structure. Not just a sequence: a solved or modeled three-dimensional arrangement of atoms, with coordinates, so that you can compute how a substitution disrupts packing interactions, changes hydrogen bond geometry, or shifts the hydrophobic core. For years, structure-based energy functions were essentially the only game in town for predicting whether a mutation stabilizes or destabilizes a protein.
That assumption has started to crack. At Scala, our stability model takes an amino acid sequence as its only required input and returns a predicted delta-Tm: the expected shift in melting temperature relative to the wild-type sequence. No structure required, no homology modeling step, no AlphaFold-predicted model as a prerequisite. This post describes how that works, where the signal actually comes from, and where the approach still has real limitations.
Why structure-free prediction is worth pursuing
Most protein engineering projects start with a sequence, not a structure. You have a FASTA file, maybe a reference organism, and a set of residue positions you want to explore. Getting a high-quality structure typically means either pulling a PDB entry (if a close homolog exists), running AlphaFold2 or ESMFold (which adds a modeling step and its own uncertainty), or waiting for a solved crystal structure (which requires the protein to express and crystallize, often only after optimization).
Each of those paths adds time and introduces its own errors. A homology model at 40% sequence identity may place a sidechain in the wrong rotameric state. An AlphaFold prediction that scores pLDDT 75 in a loop region is telling you it is uncertain about that region, which is precisely the region you may want to mutate. When the goal is to screen hundreds of variants, having structure as a prerequisite is a meaningful bottleneck.
Sequence-based prediction sidesteps this entirely. The tradeoff is that you give up the direct spatial reasoning that lets physics-based methods catch obvious steric clashes. What you gain is coverage: you can score any sequence without stopping to build a model first.
Where the signal comes from without structure
There are two main sources of stability information in a sequence. The first is evolutionary: residues that have been conserved across many homologs are usually conserved for a reason, and positions that accept diversity tend to be more tolerant of mutation. Multiple sequence alignments of a protein family encode, in aggregate, a great deal of information about which substitutions the fold can accommodate.
The second source is learned patterns from experimental stability data. Over the past decade, large-scale deep mutational scanning experiments have generated quantitative fitness data for thousands of variants of dozens of proteins. ProThermDB, FireProt, and similar curated datasets aggregate thermodynamic measurements from the literature. When a model is trained on enough of this data with sufficient architectural capacity, it begins to pick up patterns that are not obvious from sequence conservation alone: position-specific substitution effects, context-dependent tolerance windows, the influence of secondary structure propensity on stability outcomes.
Our model combines evolutionary features derived from deep multiple sequence alignments with learned embeddings from a pre-trained protein language model, calibrated against a proprietary curated stability dataset that we have extended through our early-access pilot program. The output is a delta-Tm value in degrees Celsius, with an associated confidence interval that reflects alignment depth and model certainty at that position.
A concrete example: xylanase surface engineering
One of the protein families in our pilot program was a GH11 xylanase used in a paper pulp processing application. The engineering goal was to improve stability at 65 to 70 degrees Celsius without compromising activity at pH 5. The starting variant was already reasonably stable by industrial standards (Tm around 60 degrees Celsius), but the target operating window required a further 8 to 10 degree shift.
The team had 14 surface-exposed residue positions they wanted to explore, based on prior structural knowledge. Rather than ordering all single-point variants at 14 positions (19 alternatives each = 266 variants), they ran a Scala scan first. The output flagged six positions as having high predicted positive delta-Tm potential, with specific substitutions in each case. Three of those positions had been on the team's radar; three were not. The scan also flagged two positions on the original list as having consistently neutral or slightly negative predictions, suggesting they were not worth the synthesis cost.
The team synthesized 22 variants from the top tier of the ranked output, expressed them, and measured Tm directly. Of the eight variants ranked in the top decile by predicted delta-Tm, six showed measurable positive shifts ranging from 2.3 to 7.1 degrees. The two misses were at the margins of the confident prediction zone, where alignment depth for that particular subfamily was thinner than average.
What delta-Tm is, and what it is not
It is worth being precise about what the model outputs. Delta-Tm is a predicted change in melting temperature relative to the wild-type reference, not an absolute predicted Tm value. Absolute Tm prediction from sequence alone is substantially harder because it depends on buffer conditions, protein concentration, quaternary structure state, and other factors that are not captured in sequence. The delta formulation is more tractable because many of those context factors cancel out in the comparison between a variant and its reference.
We are not claiming that a predicted delta-Tm of plus 4.2 degrees means your variant will be exactly 4.2 degrees more stable than wild type under your assay conditions. What it means is that the model assigns this variant a substantially higher probability of improved stability than a variant scoring plus 0.3 degrees or minus 1.8 degrees. The value is most informative as a ranking signal across a panel of variants, not as an absolute thermodynamic quantity.
This is an important distinction for wet-lab integration. Using the output to rank a panel of 50 variants and select the top 8 for synthesis is a well-matched use case. Using it to replace a DSF or nano-DSC measurement is not what the tool is for.
The limits of sequence-only prediction
The main failure modes of our current approach are predictable and worth knowing upfront. First, prediction accuracy degrades as alignment depth decreases. For proteins from very sparse families (fewer than a few hundred high-quality homologs in the alignment), the evolutionary signal is weak, and the model's confidence interval widens accordingly. The output is still directionally useful, but the top-1 ranking precision is lower.
Second, long-range allosteric interactions that are mediated through specific tertiary contacts are harder to capture from sequence alone. A substitution at a surface position that changes a buried cavity geometry three turns of a helix away is more likely to be mispredicted without structural context. Our hybrid model partially compensates through learned co-evolution signals in the language model embedding, but it is not a complete substitute for explicit spatial modeling in those cases.
Third, insertions and deletions are outside the current scope. The model takes a fixed-length sequence aligned to a reference; predicting the stability impact of length changes requires a different treatment of structural context that we have not yet incorporated into the main pipeline.
For engineering campaigns where the target protein has a well-populated protein family and the mutations are single-point substitutions or small combinatorial sets, the sequence-only approach performs well enough to be the first-pass filter. For membrane proteins, intrinsically disordered regions, or proteins with no close homologs, we recommend treating the output with more caution and integrating structure-based checks for the top candidates.
What this means for how you set up a campaign
The practical implication is that you can run a stability scan before any structural work is done. In a standard early-stage engineering campaign, this changes the order of operations: you see the predicted stability landscape first, use it to identify which positions and residues to prioritize, and then do structural analysis (if needed) only for the top candidates before synthesis. The sequence-scan output narrows the field substantially, which is where most of the value comes from.
We have found that teams get the most from the output when they treat it as a prioritization layer rather than a decision-maker. The scan tells you where to look harder. Whether a specific candidate goes into your synthesis list still benefits from your own mechanistic judgment about what you know structurally and functionally about the protein. That combination tends to produce better hit rates than either source of information alone.