Hybrid physics and machine learning, grounded in experimental data
Scala's stability model combines residue-level energy functions with sequence-fitness representations trained on deep mutational scanning datasets. Neither approach alone achieves the accuracy we need for practical wet-lab guidance.
Why a hybrid model outperforms either component alone
Physics-based energy functions capture the thermodynamic intuition behind stability: packing, electrostatics, solvation, and backbone strain. But they struggle on remote homologs and multi-site combinations where context matters beyond the local interaction graph.
Sequence-fitness models trained on deep mutational scanning data capture evolutionary and functional context. But they can be overfit to the training protein family and miss non-conservative substitutions outside the training distribution.
Scala's ensemble combines residue-level energy deltas with learned fitness features, weighting each contribution by a per-family confidence estimate derived from sequence identity to known training data. This gives the model calibration: it knows when to lean on physics and when to defer to sequence context.
Independent benchmark results
Evaluated on held-out test sets from ProThermDB and FireProtDB, using only sequence as input. Spearman rank correlation with experimental Tm shift or dG change.
Spearman r on ProThermDB held-out set, 1,240 single-point variants
Lipases, cellulases, xylanases subset; same evaluation protocol
Fraction of actual top-10% stabilizing variants found in model's top-10% predictions
Families with less than 20% sequence identity to any training protein
Benchmarks use publicly available datasets (ProThermDB v4, FireProtDB 2024). Full evaluation protocol available on request.
Foundation literature and Scala technical notes
Global analysis of protein folding using massively parallel design, synthesis, and testing
Science, 357(6347):168-175, 2017. DOI: 10.1126/science.aan0693
Foundational thermostability dataset used in model trainingLanguage models enable zero-shot prediction of the effects of mutations on protein function
NeurIPS 2021 (Advances in Neural Information Processing Systems, vol. 34). bioRxiv: 2021.07.09.450648
Sequence-fitness representation approach informing Scala's learned featuresScala internal technical note: confidence calibration for out-of-distribution protein families
Internal document. Available to research partners on request.
Scala internal technical noteQuestions about our methods? Talk to the team.
We are happy to discuss benchmark methodology, dataset coverage, or how the model handles your specific protein family.