The fitness landscape concept, introduced by Sewall Wright in population genetics and later adapted for protein evolution, remains one of the most productive mental models in protein engineering. The idea is simple: imagine a high-dimensional space where each axis represents a sequence position and each possible sequence corresponds to a point in that space with an associated fitness value. Engineering a protein for improved stability is then a navigation problem: find a path through sequence space that moves toward a higher-fitness region.
The challenge is that the actual landscape is not smooth, convex, or easily navigable. It has local maxima, ridges, valleys, and vast flat regions. Directed evolution's iterative mutation-and-selection cycle is an explicit acknowledgment of this: you can only take steps you can see, and you can only see one step ahead at a time. Fitness landscape methods try to change that by mapping more of the terrain before you commit to a path.
Empirical fitness landscapes from deep mutational scanning
The most detailed empirical fitness landscapes we have come from deep mutational scanning (DMS) experiments. In a DMS experiment, a library of all single amino acid substitutions at every position in a target region is generated by oligonucleotide-based mutagenesis, expressed in a suitable host, and selected for a functional property. The selected population is sequenced by next-generation sequencing, and the frequency of each variant before and after selection gives a direct fitness estimate for each mutation.
DMS has been applied to proteins ranging from bacterial enzymes to antibodies to fluorescent proteins, yielding some of the most comprehensive fitness landscapes in biology. The datasets from these experiments are not just useful for the specific protein studied. They are training resources that reveal general principles: which secondary structure contexts tolerate substitution most readily, how conservation correlates with tolerance, how the sign of individual mutation effects distributes across a protein family.
The limitation of DMS for protein stability engineering in particular is that the fitness proxy is not always thermostability. Many DMS experiments use cell growth, binding affinity, or enzyme activity as the selection signal. These correlate with stability in many cases, but the correlation is imperfect. A protein that is more thermostable but less active will be penalized in an activity-based selection regardless of its stability improvement. Stability-specific DMS datasets that use thermostability as the selection criterion directly are fewer and harder to generate but more directly informative for our purposes.
Computational landscape representations: from statistical models to language models
Before experimental DMS data is available for your target protein, or when it is unavailable at all, computational landscape representations take over. The simplest is the position-specific scoring matrix (PSSM) derived from a multiple sequence alignment: each cell represents the log-odds probability of observing a given amino acid at a given position, derived from evolutionary frequency. This is a flat, independent-site model with no interaction terms.
PSSMs are still useful, particularly as a fast check for obviously deleterious substitutions at highly conserved positions. Their main limitation is the independence assumption. The PSSM assigns a score to a substitution at position A without knowing what residue occupies position B, which is correlated with A through the protein structure. Direct coupling analysis (DCA) and related methods address this by modeling the statistical dependencies between position pairs, extracting a residue-residue contact map and pairwise potentials from the alignment. The resulting probabilistic model is a better representation of the fitness landscape for multi-site mutations than a PSSM.
Protein language models represent the current frontier of sequence-based fitness landscape modeling. Trained on hundreds of millions of protein sequences from UniRef, these models develop internal representations that capture deep evolutionary statistics across the sequence space. Predicting the fitness of a sequence variant in this framework is equivalent to asking how "likely" the model considers that variant to be in the broader space of functional protein sequences. This is not thermostability directly, but it correlates with stability strongly enough to be a useful prior for stability prediction when calibrated against experimental data.
Navigating a predicted landscape: search strategies
Having a fitness landscape representation, whether empirical or computational, changes the engineering strategy. Instead of picking mutations based on local structural analysis and trying them one by one, you can use the landscape representation to design a search strategy.
The simplest strategy is greedy hill-climbing: start from wild type, evaluate all reachable single-step variants, take the best step, repeat. This works well when the landscape is smooth near the starting sequence, but fails when the local maximum is suboptimal and reaching the global maximum requires passing through a fitness valley. Greedy hill-climbing is also what most traditional iterative engineering campaigns do, just with a much smaller view of the landscape at each step.
Broader search strategies include simulated annealing (accepting occasional downhill steps to escape local maxima), genetic algorithm approaches (maintaining a population of variants and combining them through selection and recombination), and Bayesian optimization (using a surrogate model of the landscape to suggest the next most informative experiments to run, trading off between exploration and exploitation).
We use Bayesian optimization logic in how we suggest multi-round campaign designs. If a team has completed a first round of wet-lab validation on a Scala-recommended shortlist and has experimental Tm data for those variants, we can incorporate that data as direct calibration for the target protein and suggest a second-round shortlist that updates the landscape model based on actual measurements. The second round is then targeted at regions of sequence space that the first round narrowed but did not exhaustively sample.
Ruggedness and what it means for engineering confidence
Fitness landscape ruggedness, the degree to which nearby sequences have uncorrelated fitness values, is one of the most practically important landscape properties for protein engineers. A smooth (low-ruggedness) landscape means that local search strategies work well and that positive mutations tend to be robustly positive across nearby backgrounds. A rugged (high-ruggedness) landscape means that even modest sequence changes can flip the sign of a mutation's effect.
The available evidence suggests that protein stability landscapes have intermediate ruggedness. Local regions of sequence space tend to be smoother than the global average, meaning that for a starting sequence well within a stable fold, the probability that a stabilizing mutation stays positive when combined with another nearby mutation is higher than for completely random pairs. But the landscape is not smooth enough to make the additive assumption reliable for all combinations, as epistasis data repeatedly confirms.
Practically, this means the confidence of any stability prediction decreases as you move further from the starting sequence, whether by combining more mutations or by considering more radical substitutions. A prediction for the fitness of a 5-mutation combination of a protein that has never been studied computationally is less reliable than a prediction for a single mutation in a well-characterized family. Communicating this gradient of reliability, rather than a single accuracy number that applies uniformly, is one of the design goals we have built into the Scala output format.
Landscape visualization as a communication tool
One underappreciated function of fitness landscape representations is as a communication tool within a team. When you can show a visual map of predicted stability across your target region, showing which positions are high-leverage and which are flat, the team can have a substantive conversation about where to focus their attention. "Here are the three peaks, here is the predicted interaction between peak A and peak B, here is why peak C is flagged as uncertain" is a more structured conversation than "let's try these 15 variants and see what happens."
The heatmap output from Scala's mutation scanner is designed for exactly this use: a two-dimensional view of predicted stability with positions on one axis and substitution types on the other, colored by predicted delta-Tm. It is not a dimensionally accurate representation of the actual high-dimensional sequence space, but it captures the most decision-relevant structure of the landscape for the positions you are interested in. For most engineering campaigns, that is the level of resolution you actually need for the synthesis planning conversation.