The engineering goal for most industrial enzyme applications sounds deceptively simple: keep the protein stable at high temperature while keeping it soluble at the concentrations you need to run your process. In practice, these two properties are frequently in tension, and the tradeoff between them is one of the most common sources of frustration in industrial enzyme development campaigns.
This tension has a thermodynamic basis. Many substitutions that increase the melting temperature of a protein do so by increasing core hydrophobic packing, reducing backbone flexibility, or introducing additional charged-pair interactions. Some of those same changes increase the propensity for aggregation in solution, particularly at the elevated temperatures where the stabilized protein is meant to operate. The variant that passes your DSF thermostability assay may aggregate in the process reactor at the concentration and pH your reaction requires.
Why the two properties resist simultaneous optimization
Thermostability in a folded protein is primarily governed by the stability of the folded state relative to the unfolded ensemble: you want the folded free energy minimum to be deeper and the unfolded state to be higher in energy. Solubility is a different problem: it is governed by the interactions between folded protein molecules in solution, not the folded-to-unfolded thermodynamic gap. A protein with a higher Tm can aggregate more readily if stabilizing mutations expose hydrophobic patches or reduce the surface charge density that keeps molecules from clustering.
The mechanistic conflict is at the surface. Mutations that tighten the hydrophobic core typically do not cause aggregation directly, but mutations that introduce surface-exposed hydrophobic contacts (sometimes done to stabilize a loop or a flexible linker region) can create aggregation nucleation sites. Disulfide bridges, which are among the most reliable thermostabilizing strategies for extracellular enzymes, can sometimes shift the oligomerization state in solution if a cysteine is poorly placed.
Consensus sequence approaches, which are popular for industrial enzyme stabilization because of their relative simplicity and success record, have a known failure mode: consensus residues at buried positions reliably improve Tm, but consensus residues at exposed positions can increase aggregation propensity if the consensus was derived from a set of organisms that operate at different optimal salt or pH conditions.
How Scala's model handles the multi-objective problem
Our primary stability model predicts delta-Tm, a single scalar output. It does not directly predict aggregation propensity or in-solution solubility. This is an honest statement of what the model does, not a limitation we are glossing over.
What the model does provide that is useful for solubility management is a secondary annotation on each predicted variant: an estimated change in surface hydrophobicity and a surface charge density shift relative to wild type. These are not predictions of aggregation per se. They are sequence-derived features that correlate with aggregation propensity in a wide range of protein families, and they flag variants that are at risk of introducing solubility problems even while improving thermostability.
In practice, this annotation lets teams apply a simple filter: exclude from the synthesis shortlist any variant that shows strong predicted positive delta-Tm but also shows a large predicted increase in surface hydrophobicity or a large shift in net surface charge away from the target pH optimum. This does not guarantee the filtered candidates will behave well in solution, but it removes the most obviously problematic cases before synthesis.
A case from a pilot run: carbohydrate-active enzyme stabilization
One of our pilot cohorts was working on a GH5 endoglucanase intended for use in a pre-treatment process at pH 5.5 and 60 degrees Celsius. The wild-type enzyme showed adequate activity but began losing significant activity after 30 minutes at operating temperature. The goal was to extend half-life at operating conditions without compromising activity against crystalline cellulose substrates.
The initial Scala scan identified 12 high-confidence positive delta-Tm candidates. When the surface hydrophobicity annotation was applied, three of the 12 showed predicted increases in surface hydrophobic patch area that were flagged as moderate risk for aggregation at concentrations above 5 mg/mL. Those three were moved to a lower-priority tier in the synthesis plan.
The team synthesized 9 variants from the higher-priority shortlist. Six showed improved DSF Tm of 2 to 6 degrees relative to wild type. Of those six, one showed reduced solubility at 10 mg/mL in the test buffer, confirming that the surface hydrophobicity annotation does not catch everything. The other five were soluble at working concentrations and showed improved operational stability in the half-life assay. Two of the three de-prioritized variants were also synthesized for comparison: one showed aggregation at 8 mg/mL, consistent with the flag; the other was soluble but showed reduced activity, possibly due to a loop stabilization substitution affecting substrate binding.
Practical guidance for running stability-solubility optimization campaigns
Based on this and several similar pilot runs, our current recommendation for industrial enzyme campaigns where both properties matter is the following sequence. Start with the full thermostability landscape scan to identify all positive delta-Tm candidates. Apply the surface hydrophobicity and charge density filters to remove the most at-risk variants. If you have a secondary solubility prediction tool (CamSol, Aggrescan, or your own empirical dataset from prior campaigns on the same family), apply it to the remaining candidates as a second filter pass. Take the intersection of all three filter outputs to your final synthesis list.
The result is a smaller shortlist that has been evaluated against multiple fitness dimensions before synthesis. It will not be free of surprises. Protein behavior in solution at process-relevant conditions involves factors that no current model captures completely, including the effects of other components in the reaction mixture, chelating agents, surfactants, and the specific isoform of the substrate. But the pre-synthesis filtering does reduce the number of variants that fail specifically because of a predictable and avoidable property conflict.
What the field still needs here
Honest assessment: the computational tools for solubility prediction are currently behind the tools for thermostability prediction. Experimental aggregation data at scale, equivalent to the deep mutational scanning datasets that have driven advances in stability prediction, does not yet exist in comparable volume. The aggregation and solubility annotations we provide are useful flags, not quantitative predictions with the same confidence as the delta-Tm output.
Building better solubility prediction requires either large-scale experimental aggregation datasets or better physics-based models of intermolecular interaction in solution. Both are active areas of academic and industrial research. Until those resources mature, the right practice is to use the computational tools as a filter layer with known limitations, and to not rely on them as a substitute for small-scale in-solution behavior testing of your shortlisted candidates before committing to full characterization.