When we started building the Scala platform, we had a technical team with machine learning and computational biology backgrounds and a clear view of what the model could do. What we did not have initially, and had to develop over time through early conversations and pilot engagements, was a clear picture of how a protein engineer actually uses predictions during a campaign planning session. These two things, what the model can do and how it fits into a real workflow, turn out to be quite different, and the product design decisions that came from understanding that gap shaped almost everything about what the platform looks like today.
This post is a reflection on some of those decisions: what we got wrong at first, what we changed, and why some of the choices we made are less obvious than they appear in the final product.
The first design mistake: exposing model hyperparameters
Our first interface prototype let users select from different prediction modes with different underlying configurations: a fast mode that used sequence features only, a full mode that added structural energy features when a PDB file was provided, and an experimental mode that allowed users to set alignment depth thresholds and pick which protein language model embedding layer to use. We thought this was a feature. We had users who might want to compare the two modes, or who might have opinions about alignment filtering based on their knowledge of their protein family.
What we observed in early testing was that these choices created decision paralysis rather than flexibility. Protein engineers are experts in their proteins, not in alignment filtering strategy. When we presented a choice between running with or without structural features, they did not know which would be more accurate for their specific case, and neither did we in a way we could communicate simply, because the answer was genuinely protein-family-dependent. The choice was not information to the user: it was an unresolved modeling question pushed onto them.
We removed the mode selection entirely. The platform now runs the best-available configuration for each input automatically, using the structural features if a PDB is provided and sequence-only features if not, with no option to override. Users lost flexibility they were not using and did not have good basis to use. What they gained was a simpler input form and confidence that the prediction they received was the best the system could provide for their input, not one of several possible predictions they had to choose among.
Learning what "confidence" means to a wet-lab scientist
The confidence tier system (Tier 1, 2, 3) went through multiple iterations before arriving at the current design. Our original confidence output was a continuous probability score between 0 and 1. We thought this was appropriate: it preserved the full information in the model's uncertainty estimate and let users set their own threshold based on their risk tolerance for a synthesis decision.
What we learned from pilot users was that a continuous score in the range of 0.62 to 0.79 is not interpretable when you are deciding whether to include a variant in a synthesis order. The numerical range conveys false precision, because the score itself is an estimate of uncertainty rather than a direct measurement. More practically, protein engineers think in terms of synthesis batches: they are deciding which 15 variants go in the next order, and they need a decision-ready categorization, not a fine-grained score to translate into a decision.
The three-tier system maps to three synthesis planning behaviors: Tier 1 candidates get included in the synthesis order with high confidence (you trust the prediction enough to commit the synthesis cost). Tier 2 candidates get included if synthesis budget allows and you want to hedge. Tier 3 candidates do not get included unless you have a specific structural or biological reason to override the model's low confidence signal. This is the actual decision logic that a campaign planner runs, and the tier system makes it explicit rather than requiring a translation step from score to action.
The heatmap as a primary interface object
The decision to make the mutation heatmap the primary output visualization was not obvious from a software design perspective. A ranked list with sortable columns would have been easier to build and arguably more information-dense. Heatmaps require the user to learn a visual encoding (position on x-axis, substitution type on y-axis, color for effect magnitude) before they can extract information.
We made the heatmap primary because of how protein engineers think about mutation space. When a structural biologist or biochemist looks at a set of residues and asks "what can I do here?", they think spatially along the sequence, not as a ranked list of individual candidates. The heatmap preserves the spatial structure of the sequence: position 47 is next to position 48 and 49, and you can see at a glance that position 47 has multiple blue cells (negative predictions) while position 48 has several teal cells (positive predictions). That spatial pattern is informative in a way that a ranked list does not communicate.
The pattern along the amino acid axis (y-axis) is also informative structurally. If the positive predictions at a given position are concentrated in the small hydrophobic amino acids (Ala, Val, Ile, Leu), that pattern suggests a core-packing improvement mechanism: the wild-type residue is too large for the local cavity, and smaller substitutions relieve the packing strain. This mechanistic inference happens visually from the heatmap in seconds; extracting it from a ranked list would require cross-referencing the amino acid types manually.
What we did not build: variant design recommendations
We chose not to include an automated variant design recommendation feature that suggests "make this specific set of mutations in this specific combination." Users asked for this feature in early interviews, and it would have been straightforward to implement as a simple ranking-and-filtering step on top of the existing prediction output.
The reason we held back is about epistemic responsibility. A recommendation that says "make these 3 mutations" requires the system to be confident not just in the individual stability predictions but in the engineering strategy: which positions to prioritize, how many mutations to combine in one round, whether to run one large panel or two smaller sequential rounds. These decisions depend on the team's synthesis budget, timeline constraints, what assay they are using to characterize variants, and how confident they are in the model's predictions for their specific protein family.
All of those factors are context the system does not have. A recommendation built without that context would be confident-sounding but not actually well-grounded, and in protein engineering, confident recommendations that turn out to be wrong have costs: synthesis resources committed, characterization time spent, and engineering cycles burned on a direction that the team would not have taken if the tool had been more honest about the limits of its judgment.
Instead, we surface the information the recommendation would have been based on, and let the engineer make the synthesis planning decision with that information. The heatmap, tier annotations, and epistasis flags are all inputs to that decision. The engineer provides the missing context from their knowledge of their process, their budget, and their protein. The combination of computational triage output plus domain expert judgment produces better synthesis plans than the computational output alone could produce.
The documentation challenge: explaining what the model is predicting
One of the ongoing product challenges is communicating what the model is and is not predicting. The output is labeled delta-Tm, which is technically accurate: the model predicts the change in thermal denaturation midpoint. But several important caveats attach to that label that are easy to miss if you are in a hurry.
First, the prediction is for the isolated protein under standard buffer conditions. Process conditions often differ from standard buffer conditions in ways that affect thermal denaturation: different pH, the presence of co-solvents, substrates, or cofactors. The Tm under process conditions may be meaningfully different from the predicted Tm under standard conditions, and the effect of mutations on process-condition Tm may not track exactly with the effect on standard-condition Tm.
Second, Tm is not the same as functional stability. A mutation that raises the thermal denaturation midpoint may also reduce the rate of catalysis at the operating temperature, because it reduces the backbone flexibility that enzyme turnover requires. Higher Tm is a necessary but not sufficient condition for improved process performance.
We address this with documentation and in-interface tooltip text, but the honest assessment is that a tooltip is not the right format for a nuance this important. We are working on a more prominent pre-scan checklist that surfaces these caveats as part of the workflow before a user interprets results, rather than as background documentation they may or may not read. The goal is not to discourage use: the predictions are useful even given these caveats. The goal is informed use, where the engineer's interpretation of the output is grounded in a correct understanding of what the number means and does not mean.
Where we are now
The platform in its current state reflects about 18 months of iteration on these design questions. The core prediction capability has not changed dramatically since we first built it: the model architecture and the training approach are largely stable. What has changed substantially is the interface through which that capability reaches protein engineers. Every simplification we made, every removed option, every visual encoding choice, came from a specific conversation with someone who was trying to use the tool to make a real synthesis decision and running into friction that we could eliminate without sacrificing the information they actually needed.
The ongoing challenge is that the user population is not uniform. A structural biologist with computational experience reads the heatmap differently from a biochemist whose primary expertise is assay development and enzyme kinetics. Designing an interface that is not condescending to the expert while also being usable to someone who has never seen a fitness landscape visualization before is an unsolved problem. We have made choices that favor the less computationally experienced user at some cost to the expert user, and there are cases where that tradeoff is wrong. We expect to keep revising it.