We are cautious about publishing results from pilot runs before we are confident enough in what they actually show. This case study is an account of one of our first extended pilot engagements, written with the pilot partner's agreement to share non-identifying technical details. The purpose is not to demonstrate that our model is perfect. It is to show what the output looked like, where it was right, where it was wrong, and what the team actually learned from the process.
The pilot involved a bacterial lipase from the GDSL family, used in a biodiesel conversion process where operating temperature was a primary engineering constraint. The team had been working with this enzyme for about two years prior to our engagement. They had prior mutagenesis data on about 40 single-point variants from an earlier campaign and a solved crystal structure at 1.9 angstrom resolution. This was a relatively favorable situation for our model: good structural information, an available prior dataset for calibration context, and a well-represented enzyme family in public databases.
Setting up the scan
The team identified 18 target positions based on two criteria: residues within 12 angstroms of the active site serine that the prior campaign had not covered, and surface-exposed positions where previous literature suggested stability improvement potential in related GDSL lipases. Their original plan was to synthesize all 18 positions exhaustively, all 19 possible amino acids per position: 342 variants. Budget and timeline constraints were the limiting factors, not experimental capacity.
They ran the Scala scan against those 18 positions using their crystal structure as the input structure and the reference wild-type sequence. The scan returned predicted delta-Tm values for all single-point substitutions plus predicted pairwise combinations for the top 6 individual-position hits. The output also included confidence tier annotations and the surface hydrophobicity change estimate for each variant.
The first thing the team noticed was that the scan identified only 6 positions as having high-confidence predicted positive effects: positions 47, 112, 156, 203, 219, and 267. The remaining 12 positions showed predominantly neutral or negative predictions. Three positions that the team had included based on a recent publication on a related lipase from a different organism showed consistent negative predictions across most substitutions. When we looked at the alignment for those positions, there was a reason: the related organism's lipase had a very different packing environment at those positions due to a different secondary structure element in that region. The published mutation had worked for a different structural context.
What got synthesized and what happened
After reviewing the scan output with their structural knowledge, the team decided to synthesize variants at 11 of the 18 positions rather than all 18, and at those 11 positions they synthesized only the top 2 predicted substitutions per position based on the Scala ranking, plus the wild-type reference. Total synthesis: 23 variants plus wild type. This was compared to the original plan of 342 variants: an 85% reduction in synthesis scope.
Expression and purification proceeded in an E. coli system the team was already using. One variant (at position 47) failed to express at useful levels. The remaining 22 were expressed, purified to acceptable homogeneity by SEC, and characterized by DSF in triplicate.
Results summary:
- Of the 6 predicted high-confidence positions, variants at 5 showed positive Tm shifts in the wet lab. The 6th (position 156) showed negligible shifts across both substitutions tested.
- The largest positive shifts came from position 219 (Ile to Val, plus 4.8 degrees Celsius) and position 203 (Asp to Asn, plus 3.9 degrees Celsius). Both had been in the top tier of the computational ranking.
- Of the 5 variants synthesized at positions the scan had predicted as predominantly neutral, all showed Tm shifts between minus 0.4 and plus 1.1 degrees Celsius. The distribution was consistent with what the model had predicted as effectively flat.
- The double mutant combining positions 219 and 203 showed a Tm improvement of 7.6 degrees Celsius, close to the additive sum (8.7 degrees), indicating limited negative epistasis between those positions. The scan had flagged this pair as low co-evolutionary coupling, consistent with approximately additive behavior.
Where the model missed
Position 156 was a clear miss. The model predicted positive effects for Ala and Gly substitutions at that position, both showing predicted delta-Tm of plus 2.1 and plus 1.9 degrees respectively. The wet-lab data showed no significant improvement: Ala showed plus 0.3 degrees (within measurement noise) and Gly showed minus 0.6 degrees. When we went back to investigate, the likely cause was that position 156 is at the edge of a loop that adopts a conformation in the crystal structure that differs from the ensemble-average solution behavior. The static backbone energy calculation had overweighted a conformation that was not representative of the actual solution state at operating temperature.
This is a known failure mode for fixed-backbone structure-based predictions near flexible loops, and it was present even in our hybrid model. The position had scored as high confidence because the alignment depth was good and the evolutionary features were favorable. The structural information was misleading because the crystal form captured a locally low-energy conformation that is not the operative state in solution. We have since added a loop flexibility indicator to the confidence tier annotation that flags positions where the crystal structure's B-factors suggest local flexibility above a threshold.
A second category of misprediction: two variants at positions 47 and 112 that the scan predicted as moderate positive effects showed neutral outcomes in DSF. In both cases, the wet-lab measurement was within the noise range of the assay, and the prediction uncertainty interval did overlap with zero. These are technically within the model's stated uncertainty; the team had taken the central estimate as more informative than the confidence interval. This is a calibration communication issue on our side, and we have since changed how we present the output for predictions that fall near the neutral zone.
The practical outcome for the campaign
The pilot partner combined the two best-performing single substitutions (positions 219 and 203) and proceeded with functional characterization of the double mutant. The combined variant showed the expected improvement in Tm and retained greater than 90% of wild-type lipase activity against their process substrate. Half-life at 60 degrees Celsius in the process buffer was extended from approximately 3.5 hours for wild type to approximately 11 hours for the double mutant, which met their process target.
That outcome required testing 23 variants instead of the original 342. The time saved at the synthesis and screening stage was approximately 6 weeks in the team's estimate, accounting for the cloning, expression, and characterization pipeline they normally ran. Whether that represents a typical result for Scala at this stage of development is hard to say from a single case. But it is representative of the specific value case we are trying to make: knowing where to look before you synthesize changes the economics of the entire campaign.
What this case does and does not prove
One pilot case with favorable results is not a validation of the model at scale. The lipase family is well-represented in the training data, the crystal structure was high quality, and the engineering goal was narrowly defined. A less favorable combination of those factors would likely produce different results. We are publishing this case because it is an honest record of a real engagement: the hits, the misses, and the follow-up analysis.
What it does show is that the output was decision-useful. The shortlist contained the actual high performers, the missed candidates were identifiable in retrospect as falling in known failure mode categories, and the overall campaign timeline was meaningfully compressed. That is the operating criterion we are optimizing against as we continue to develop the model.