Inference of experimental protein fitness landscapes
Probabilistic models that infer sequence-function relationships directly from high-throughput selection experiments
High-throughput selection experiments such as deep mutational scanning, phage display and directed evolution probe millions of sequence variants in a single assay, and are now the main source of quantitative sequence-function data. What they return, however, is not a measurement of function but a set of read counts across selection rounds, shaped by the library composition, the stringency of each round and the amplification steps in between. Models trained directly on enrichment ratios inherit these artifacts and conflate the landscape with the experiment that sampled it. Evolutionary models, at the other end, capture the constraints acting on natural sequences but are blind to the specific biochemical activity being selected for.
I develop probabilistic models that describe the selection experiment explicitly, inferring a sequence-dependent selection energy by maximum likelihood on the observed counts. Because the experimental process is part of the likelihood, the inferred landscape is separated from round-specific effects, and epistatic couplings can be estimated without assuming additivity. These models are generative: they do not only score existing variants but define a distribution from which optimized sequences can be sampled, either directly or by guiding pretrained protein generative models toward regions of high predicted activity. The framework is developed on data produced by experimental collaborators and validated prospectively on newly designed sequences.
Institutions
Collaborators
Funding
TODO