Back to Blog
product

Introducing the Scala Mutation Scanner

By Ravit Netzer

Introducing the Scala Mutation Scanner

When we started building Scala's platform, we had a single-sequence prediction interface: submit a sequence plus a variant, get a stability delta back. That interface made sense for evaluating specific hypotheses. You have a mutation you already believe in and want a computational sanity check. But it was not the right shape for the task that most protein engineers actually face, which is not "what does this one mutation do?" but "which of these 400 possible mutations should I actually make?"

The mutation scanner is our answer to that question. It maps stability predictions across your full variant space in a single job, outputs a ranked candidate list, and visualizes the results as an interactive heatmap. This post describes what it does, how to set it up, and where the current version falls short (which is worth knowing before you design a campaign around it).

What the scanner does

The scanner takes a target protein sequence and a set of positions you want to explore. It then evaluates all single amino acid substitutions at each specified position, producing a prediction matrix covering every possible single-point change in your target region. For a 15-position scan, that is 285 predictions (15 positions times 19 non-wild-type amino acids). For a 20-position scan, 380 predictions.

In addition to single-point predictions, the scanner runs pairwise combination scoring for all pairs of positions that showed top-tier individual effects. If your single-point scan identifies 6 high-confidence stabilizing positions, the scanner automatically computes predictions for all 15 pairwise combinations of those 6 positions, using the highest-scoring substitution at each position as the combination input. This catches the most likely synergistic pairs and flags predicted epistatic conflicts before you commit synthesis resources to combinations that are unlikely to work.

The output is a CSV download containing the full prediction matrix and a confidence annotation for each entry, plus an interactive heatmap visualization in the platform UI. The heatmap lets you filter by confidence tier, sort by predicted delta-Tm, and export a synthesis shortlist of your selected candidates directly from the interface.

How to set up a scan

Setup requires three inputs: a protein sequence in FASTA format, a list of residue positions to scan (specified as a comma-separated list of position numbers or a range), and optionally a structural context file (PDB format) that the model can use to supplement the sequence-based predictions with position-specific structural annotations.

The structural context input is optional. If you provide it, the model uses it to annotate each position in the scan with secondary structure assignment, accessible surface area, and local packing density. These annotations do not directly change the stability predictions; they are auxiliary information that helps you interpret the results. A position flagged as highly buried with low accessible surface area and high local packing density is in a context where mutations are structurally constrained, which informs how much latitude you have in exploring that position.

If you do not have a structure, the scanner runs on sequence features alone and still produces valid predictions, with the caveat that the position context annotations will be absent and the confidence tier assignments will be based on sequence and alignment features only. For proteins with high alignment depth, this works well. For sparse families, providing a structure (even a computationally predicted one from AlphaFold2 or ESMFold, with appropriate caveats about loop accuracy) improves the tier annotation quality.

Reading the heatmap output

The heatmap displays positions on the x-axis and amino acid substitutions on the y-axis. Cell color encodes predicted delta-Tm: bright teal for predicted positive effects, neutral gray for near-zero effects, and dark blue for predicted negative effects. The intensity scales with the magnitude of the predicted effect.

Several visual patterns are immediately informative. Columns that are predominantly gray or dark blue across all substitutions identify positions where the model predicts that most changes are neutral or destabilizing: these are likely structurally constrained positions where you have limited engineering leverage. Columns with clusters of bright teal cells identify high-leverage positions where several substitutions are predicted to improve stability. Bright teal clusters concentrated in one region of the amino acid axis (e.g., all the small hydrophobic amino acids in a buried position) are mechanistically interpretable as core-packing improvements.

The pairwise combination panel sits below the main single-point heatmap. Each cell in this panel shows the predicted delta-Tm for the double mutant combination of the top-ranked substitution at the row position and the top-ranked substitution at the column position. Cells are color-annotated with a co-evolutionary coupling indicator that flags pairs likely to have non-additive effects.

Interpreting the confidence tiers in the heatmap

Each cell in the heatmap carries one of three confidence tier indicators, displayed as a small badge in the corner of the cell. Tier 1 predictions have adequate alignment depth, consistent co-evolutionary support, and predicted magnitude above the noise floor. These are the candidates you can take seriously for synthesis planning. Tier 2 predictions have acceptable alignment support but more prediction uncertainty: the confidence interval is wider, and the central estimate should be interpreted as directional rather than precise. Tier 3 predictions are in low-confidence territory: alignment depth is inadequate, co-evolutionary signal is weak, or the predicted effect is within the noise range of the model.

A common mistake we see in early scanner users: filtering the heatmap by predicted delta-Tm alone and ignoring the tier column. A Tier 3 prediction showing plus 4 degrees Celsius is not a strong positive candidate. It is a prediction made with low confidence that happens to land on a high value. If you synthesize based on magnitude alone without filtering by tier, you will find that your hit rate on Tier 3 candidates is substantially lower than on Tier 1.

Current limitations

The scanner in its current form covers single-point and pairwise combinations at the positions you specify. Triple and higher-order combinations are not in scope for the automated scan. The combinatorial space grows too fast for reliable predictions at that level, and the training data for higher-order combinations in most protein families is sparse. If you are trying to stack three or four mutations from the scanner's top-tier list, use the individual prediction interface to evaluate specific combinations you have reason to believe are promising rather than relying on extrapolated higher-order predictions.

Insertions and deletions are also outside the current scanner scope. We support only substitution mutations. Predicting the stability effect of length changes requires different structural modeling that we have not yet integrated into the main scan workflow.

For very large proteins (above approximately 800 residues in the current release), scan runtime increases substantially due to alignment computation time. We are working on alignment caching and partial-sequence scanning options that will help with large targets, but currently the scan is optimized for the protein sizes most common in industrial enzyme and therapeutic enzyme engineering campaigns.

What we will build next

The features we are actively developing for the next scanner release are: background-specific scanning (you specify a variant background that already carries one or more mutations, and the scan evaluates additional substitutions on that background rather than always defaulting to wild type); improved epistasis flag specificity that provides a directional prediction of the coupling effect rather than just flagging that coupling is likely; and a multi-target batch mode that lets you run the same scan parameters against multiple related sequences in a single job submission.

We will also be improving the export format to include a synthesis-ready variant ID column that maps directly to oligo ordering conventions for common synthesis vendors, which should eliminate the manual formatting step that currently sits between the output CSV and the oligo order sheet.