Building a prediction pipeline that takes a raw protein sequence and outputs a reliable stability delta score involves more design decisions than it might appear. The conceptual path is straightforward: sequence goes in, number comes out. The implementation is where choices compound and tradeoffs accumulate. This post describes the architectural decisions behind Scala's core sequence-to-stability pipeline, including why we made the choices we did and what we would do differently if we were starting from scratch today.
Stage one: sequence input and alignment construction
The first stage of the pipeline takes a FASTA-formatted amino acid sequence as input. The sequence is validated for basic formatting (single-letter codes, no ambiguity characters), length bounds (we currently support sequences between 50 and 1,200 residues), and absence of characters that would cause alignment failures downstream.
The alignment stage is the most computationally expensive part of the pipeline for novel protein families. We use HHblits to search the sequence against the UniRef90 database and the Big Fantastic Database (BFD), constructing a multiple sequence alignment (MSA) with typically a few hundred to a few thousand hits depending on the protein family. The MSA is then filtered by sequence identity (removing near-redundant sequences above 90% identity) and alignment quality (minimum alignment length and minimum number of matched columns).
One design decision we made early and do not regret: we cache alignment results for the specific query sequence and report the alignment depth as a primary quality metric visible in the output. Users who have run the same protein before see their alignment reused, which substantially speeds up repeat queries. The alignment depth number tells them immediately how much evolutionary context the model has for their protein.
A decision we might change in retrospect: our initial alignment pipeline weighted all homologs equally by default. Proteins from distantly related organisms with very different optimal growth temperatures introduce evolutionary noise if treated as equivalent context for a mesophilic target. We have since added taxonomy-aware weighting to down-weight thermophilic and psychrophilic homologs when the query is from a mesophile, but this was a retrofit rather than a founding architectural choice.
Stage two: feature extraction
From the MSA, we extract three categories of features. The first is the per-position substitution matrix: for each position in the alignment, the log-odds probability of observing each amino acid, derived from the observed frequency across homologs with pseudocount smoothing to handle sparse positions. This is the evolutionary conservation signal.
The second category is direct coupling analysis (DCA) features: for each pair of positions in the sequence, the statistical coupling strength derived from the co-variation pattern in the alignment. These features encode residue-residue dependencies that are largely invisible to single-site statistics. Computing DCA at full pairwise coverage for a 300-residue protein gives roughly 44,000 position pairs; we represent these as a sparse matrix retaining the top-ranked couplings by mutual information.
The third category is protein language model (PLM) embeddings. We pass the query sequence through a pre-trained protein language model and extract the residue-level embedding vectors from one of the middle layers. These embeddings capture learned representations of local and global sequence context in a high-dimensional continuous space. The key design decision here was which model to use and which layer to extract from. We experimented with several embedding extraction configurations and found that middle layers of a large-parameter ESM-family model gave the best correlation with our stability labels, consistent with the pattern reported in the academic literature for biological function prediction tasks.
We concatenate these three feature types into a per-position representation that is then fed into the prediction head. The concatenated representation has been chosen to be computable for any protein that has an alignment depth above our minimum threshold, without requiring a solved or predicted structure.
Stage three: the prediction head
The prediction head is a relatively lightweight neural architecture by current standards. It is a position-wise regression model with attention layers that allow information to propagate between positions before the final per-position output. For a single-point mutation, we compute the prediction for the mutated position in the context of the full wild-type sequence, extracting a delta score by subtracting the wild-type position score from the mutant position score.
One deliberate decision: we kept the prediction head simple. The feature extraction stage carries most of the representational work, and the prediction head only needs to learn a calibrated mapping from those features to stability labels. We have found that increasing the complexity of the prediction head beyond a certain point produces overfitting on the training set without improvement on prospective validation cases. This is a common pattern in low-data biological prediction problems, where the label density is sparse relative to the representation dimensionality.
The training objective is the Pearson correlation between predicted and measured delta-Tm on the training set, with a mean absolute error penalty that prevents the model from trading accuracy on easy cases for more correlated but miscalibrated predictions on hard cases. The label smoothing technique we use on the boundary region near neutral (delta-Tm between minus 1 and plus 1 degrees Celsius) was added after observing that the model had a tendency to predict very small values as reliably neutral when they were actually near-neutral with noise, which led to underconfident predictions in the moderate-effect range.
Stage four: uncertainty quantification
Every prediction outputs a central estimate and a confidence interval. The confidence interval is derived from an ensemble of predictions with different dropout configurations active during inference (MC dropout), calibrated against a held-out validation set. The confidence tier annotation (Tier 1, 2, 3) is then assigned based on two criteria: the width of the confidence interval relative to the predicted magnitude, and the alignment depth at the query position.
We initially implemented a simpler uncertainty estimate: just the prediction variance across the dropout ensemble. This was not well-calibrated against actual error: the variance was reliable in ordering predictions from more to less certain, but the absolute interval widths were systematically too narrow for novel protein families not represented in the training set. The alignment-depth-weighted calibration adjustment was added after seeing this systematically in pilot data.
Stage five: output formatting and annotation
The final stage takes the prediction values and uncertainty estimates and formats them into the output delivered to the user. The key choices here are: what to include in the download CSV, how to present the heatmap visualization, and which annotations to surface at the prediction level versus the protein level.
We settled on reporting the following per-mutation: predicted delta-Tm, lower and upper 90% confidence interval bounds, confidence tier (1, 2, or 3), predicted surface hydrophobicity change, and a flag for whether the position has high co-evolutionary coupling with any other position in the query region. At the protein level, the report includes total alignment depth, taxonomy distribution of the alignment, and the fraction of the queried positions that fall into each confidence tier.
The most contested internal decision was how prominently to surface the confidence tier. We initially placed it in a secondary column in the output CSV that users had to know to look for. After several pilot cases where teams missed the tier-3 annotations and drew over-confident conclusions from the predictions, we moved it to the primary output view and added a visual highlight for any tier-3 prediction. Precision in communication matters as much as prediction accuracy when the output is being used to make synthesis decisions.