There is a real tension in computational protein biology between the physics-based tradition and the machine learning tradition. Structural biologists and computational chemists trained in molecular mechanics are often skeptical of pure data-driven approaches: you trained a neural network on a curated database and it predicts stability, but do you know why? Machine learning practitioners working in protein science often push back: the physics-based models have known systematic errors and are computationally expensive; why impose a framework that has documented failure modes when you can learn from the data?
This tension is not merely philosophical. It affects what models get built, how they are evaluated, and what they are good for. At Scala, we chose to build a hybrid model because we believe both traditions capture something real about protein stability that the other misses, and that the failure modes of the two approaches are genuinely complementary rather than overlapping. This post describes that reasoning and the specific choices we made to integrate the two.
What structural biology contributes that learning alone does not
The physics-based tradition in protein stability computation is grounded in thermodynamics: the stability of a protein fold is the free energy difference between the folded and unfolded states, and that difference can in principle be computed from atomic-resolution structural coordinates. The Lennard-Jones potential captures van der Waals interactions, electrostatic terms capture charge-charge and dipole interactions, and solvation terms (implicit or explicit) capture the energetic cost of burying or exposing polar and nonpolar groups.
This framework has a key property that pure data-driven approaches lack: mechanistic interpretability. When a physics-based model predicts that substituting a buried leucine with an alanine destabilizes the protein, the reason is traceable: the alanine sidechain is smaller, the resulting cavity reduces the van der Waals contact energy, and the protein loses a component of its stabilization energy. You can explain the prediction in structural terms that a structural biologist will recognize as physically meaningful.
This interpretability is not just aesthetically useful. It is a quality control mechanism. If a physics-based prediction is obviously wrong, you can trace the error to a specific term: perhaps the backbone was poorly refined at that position, or the solvation model is treating an unusual environment incorrectly. The error is diagnosable in a way that the error from a neural network's hidden layer is not.
Physics-based features also perform well in a specific regime: buried core positions in well-studied protein families with high-quality structures. For a mutation that directly changes the packing interactions of a hydrophobic core residue, the physics-based calculation captures the dominant stabilization mechanism directly.
What machine learning adds that structural biology alone does not
The limitations of physics-based approaches are well-documented. The force field parameters are empirically derived from a sample of molecular systems that does not perfectly represent every protein. Electrostatics are poorly handled by most implicit solvation models at charged surface positions. Fixed backbone calculations miss the coupling between substitution and backbone rearrangement. And crucially, all structure-based approaches require a structure: coordinate files with sufficient quality to compute meaningful energetics.
Sequence-based machine learning approaches sidestep the structure requirement by learning directly from the evolutionary record. A protein language model trained on hundreds of millions of sequences learns, implicitly, how the evolutionary process has evaluated billions of natural experiments in protein design. When it assigns a log-likelihood to a sequence variant, it is drawing on a much larger evidential base than any single crystal structure or thermodynamic database could provide.
The complementary strength of learned features is at surface positions, in sparse-alignment families, and for mutation types where backbone flexibility is relevant. These are exactly the cases where physics-based methods show their largest systematic errors. A substitution at a surface-exposed charged residue, where the physics-based electrostatics model is least reliable, is a case where the evolutionary statistics from a deep alignment are more informative.
The design of our hybrid model
In designing the hybrid, we faced a choice about how to combine the two types of features: concatenate them and let the network learn the combination, or use structured fusion that preserves interpretability at the physics contribution level. We chose a structured approach: the physics-based features and the sequence-based features enter the model as separate streams that contribute independently to an intermediate representation before being jointly processed by the prediction head. This allows us to interrogate the relative weight of each stream for a given prediction.
The physics-based features we include are derived from energy function evaluations on the input structure (when available): per-residue van der Waals energy, side-chain burial estimate, hydrogen bond geometry score for positions within hydrogen bonding distance of the mutation site, and local electrostatic environment. These are relatively fast to compute and capture the structural context in a form the model can condition on.
The sequence-based features come from two sources: the position-specific statistics from the multiple sequence alignment (conservation, substitution frequencies, co-evolutionary coupling from DCA) and the residue-level embeddings from the pre-trained protein language model. The PLM embeddings are the highest-dimensional input to the model and carry the most complex representation of sequence context.
The training data that calibrates the relative weighting of these two streams comes from our curated stability dataset, which includes systematic coverage of both core and surface mutations and both well-studied and less common protein families. The dataset was assembled with the deliberate goal of maintaining this coverage balance, because a training set dominated by core mutations in well-studied families would over-learn the physics-based signal at the expense of the learned signal's contribution.
A practical consequence: what to expect at different positions
The hybrid design means that Scala's predictions have different dominant contributions depending on where in a protein the mutation is. For mutations at highly conserved buried core positions in proteins with good structures and deep alignments, the physics-based and sequence-based features tend to agree, and predictions are high confidence. For mutations at surface positions in proteins with moderate alignment depth, the sequence-based features carry more weight, and the physics-based features serve primarily as a consistency check. For mutations in flexible regions or near termini, neither stream is as reliable, and both contribute to a wider uncertainty interval.
Understanding this position-dependent behavior helps interpret the output. A high-confidence prediction at a buried core position is high confidence because two independent lines of evidence (physics and evolution) agree. A moderate-confidence prediction at a surface loop position is moderate confidence because one stream is less reliable in that structural context, not because the model has failed. The confidence tier annotation communicates this, but the underlying reason is the structural context.
Where the approach still falls short
Hybrid models have their own failure modes. When the physics-based and sequence-based streams disagree substantially, the model must adjudicate, and the adjudication is learned from training data that may not cover the specific structural context creating the conflict. Cases where the physics-based features are systematically wrong for unusual reasons (an atypical binding metal, a photoactive chromophore, a crosslink that the energy function does not model) and the sequence-based features are insufficiently informative can produce predictions that are confidently wrong rather than appropriately uncertain.
We try to identify these cases through the confidence tier system, but they are not always caught. Proteins with non-standard chemistry or unusual stabilization mechanisms are honest limitations of the current model. That applies to most current computational stability tools, and we do not claim it as unique to our approach, but it is worth naming explicitly rather than leaving users to discover it during a campaign.