When we ran the first round of Scala's early-access pilot program, our validation set was intentionally narrow. We had curated training data from published deep mutational scanning experiments and from aggregated thermodynamic databases, but our holdout test sets were drawn from the same distribution. That is a common situation for a new model: the data you trained on shapes what you can evaluate honestly.
The pilot program changed that. Each cohort of early-access partners brought their own protein families, their own assay conditions, and their own experimental workflows. Some of those protein families had sparse representation in public databases. Some used assay protocols that differed meaningfully from the standard differential scanning fluorimetry setup that dominates the published literature. The result was a more honest test of whether our model generalized, and a set of failure modes we could not have diagnosed from in-silico benchmarking alone.
The gap between benchmark accuracy and field accuracy
A model that achieves strong Pearson correlation on the FireProt benchmark or the ProThermDB test set is doing something real, but it is not the same thing as performing well on the specific protein your team is trying to stabilize. Public benchmarks are skewed toward well-studied protein families with large structural databases, primarily small soluble globular proteins that crystallize readily and have been extensively mutagenized. If your target is a membrane-associated enzyme, an intrinsically disordered region flanking a stable domain, or an enzyme from a poorly characterized microbial family, the benchmark accuracy number tells you much less than you might expect.
This is not a criticism unique to Scala's approach. It applies to essentially every stability prediction tool. The difference between field accuracy and benchmark accuracy is usually not detectable until you run your model against real engineering campaigns with real proteins and real wet-lab readouts.
The pilot program was designed specifically to surface this gap early, while we are still building the model, rather than after it has shipped to general availability.
What changed as we added more cohorts
The first two cohorts in our pilot program were both working with well-characterized bacterial enzyme families: a serine protease and a beta-glucosidase. Both are represented extensively in public stability databases. Our predictions for these proteins performed well from the start, with top-3 recall rates (the fraction of actual experimental top performers that appear in our top 3 predicted candidates) consistently above 70%.
The third cohort introduced a fungal oxidoreductase from a less-studied family. Alignment depth for this protein was significantly lower, and the published experimental data on this family was sparse. Our initial predictions on their variant panel showed notably higher error rates: several variants we predicted as neutral turned out to show moderate positive stability shifts in their DSF assay. We could trace this to the alignment layer. The model was underweighting certain position effects because it lacked the evolutionary context to recognize them.
We updated the multiple sequence alignment curation for this protein subfamily using additional sequences from NCBI that had not been in our original alignment pipeline. The recalibrated predictions on a held-out subset of their panel improved substantially. That recalibration then transferred: the same alignment update improved accuracy for a distantly related oxidoreductase that a later cohort brought in from a different organism.
The protein family coverage problem
One pattern that emerged clearly across six pilot cohorts was a correlation between alignment depth and model reliability. Proteins with more than roughly 3,000 high-quality homologs in UniRef90 at appropriate filtering thresholds showed much more consistent prediction accuracy than proteins with under 500 homologs. This was not surprising, but the steepness of the falloff was steeper than we expected.
Below about 300 effective sequences in the alignment, our confidence intervals became so wide that the rankings were only weakly informative. We are not hiding this from pilot users. The output explicitly reports the alignment depth for each query protein, and the confidence tier system flags low-depth cases before the team spends synthesis budget on candidates whose predicted improvements rest on a thin evidential base.
Getting more diverse protein families into the pilot has let us build an empirical map of where alignment depth becomes a limiting factor. This informs how we are prioritizing the next phase of training data collection. We are actively seeking pilot partners working with enzyme families in underrepresented taxonomic clades, particularly fungal and archaea-derived industrially relevant proteins, where public databases are patchier.
Assay protocol variation and what it teaches us
A less expected source of learning from the pilot cohorts was the variation in experimental protocols. Our initial training data was heavily weighted toward Tm measurements from differential scanning fluorimetry using SYPRO Orange dye. Several pilot partners used alternative methods: some used circular dichroism thermal ramp experiments, some used differential scanning calorimetry for better quantitative accuracy, and one cohort used a fluorescence-based split reporter assay as a proxy for expression-linked stability.
The correlations between our predicted delta-Tm values and the readout from DSC were generally stronger than with the DSF approximations. This was expected, since DSC measures the actual enthalpy of unfolding, while DSF is an approximation that depends on the dye's affinity for exposed hydrophobic surfaces. But understanding where that gap falls, and how large it is for different protein architectures, required seeing the actual data from partners using both methods on related variants.
This is not a model-accuracy issue per se. It is a calibration issue: knowing that a predicted delta-Tm of plus 3 degrees Celsius translates to a certain expected change in DSF apparent Tm, and a somewhat different expected change in DSC calorimetric Tm, lets us communicate what the output means more precisely to users using different assay workflows.
Closing the loop from pilot data into the model
A critical design decision for the early-access program was data sharing. Pilot partners contribute experimental measurements back to Scala's training dataset under a data sharing agreement, with appropriate confidentiality on specific protein identities and proprietary variant sequences. The stability measurements (Tm, delta-Tm values) come back to us attributed to protein family and measured positions, without revealing which specific commercial application the partner was working on.
This creates a gradual improvement cycle. Each cohort's experimental data broadens the empirical base for the next round of model calibration. The improvements are not dramatic cycle-over-cycle, but they are consistent. After six pilot cohorts, our top-3 recall rate on novel protein families not represented in our original training set is materially better than it was at program launch.
We want to be honest about the timescale here. The model is not improving in real time as you submit variants. The update cycle is quarterly: we retrain and recalibrate after each batch of pilot data is processed and quality-checked. Users who joined the pilot in its first quarter are using a model that has since been retrained with data from subsequent cohorts. That is expected behavior for an early-stage product, and we communicate the training snapshot version in the output metadata.
What this means for early-access users
If you are using Scala during the early-access program, the accuracy you see today is not the ceiling. The model is still improving as the pilot cohort data base grows. Proteins from well-characterized families will already perform well. Proteins from sparse families may show higher uncertainty, and the confidence tier output will tell you that directly.
The flip side is that pilot partners working with challenging proteins are contributing to improvement that benefits everyone in subsequent rounds. That is a genuine feedback loop, not a marketing claim about community-driven development. The data sharing arrangement is specifically structured to make this improvement cycle work while protecting each partner's proprietary information.