How Can mRNA Display Data Be Integrated with AI Peptide Design?
The most useful AI signal from mRNA display is not a final-round hit list, but a selection trajectory combining enrichment, counter-selection, sequence families, NCAA and cyclization chemistry, and experimental validation.
How Can mRNA Display Data Be Integrated with AI Peptide Design?
A tempting way to build an AI model from an mRNA display campaign is to label the most abundant sequences in the final round as “positive,” generate random sequences as “negative,” and train a classifier. This is simple, but it discards much of the experiment and can teach the model the wrong lesson.
Final-round abundance is influenced by initial representation, transcription and translation efficiency, fusion formation, peptide maturation, NCAA incorporation, cyclization, target binding, washing, recovery, PCR and sequencing. Random sequences, meanwhile, are not experimentally rejected molecules. A classifier may therefore learn library syntax, amplification bias or sequence-family identity instead of selective target binding.
The better starting point is the full experimental history. An mRNA display experiment generates a trajectory, not just a hit list. AI should learn from that trajectory, with its chemistry and selection context, and then return testable priorities to the laboratory.
What Data Does mRNA Display Actually Produce?
The useful data set has several connected layers.
| Data layer | Representative fields | Why it matters |
|---|---|---|
| Sequence | nucleotide sequence, peptide sequence, modifications, NCAA positions, cyclization anchors | Defines the encoded and intended molecular identity |
| Selection round | R0/input and R1–Rn counts, normalized abundance, enrichment and depletion | Describes how a sequence changes over time |
| Conditions | target, counter-target, matrix control, target presentation, stringency, replicate | Defines the experimental pressure behind each observation |
| Family | cluster, motif, conserved and variable positions, family expansion | Separates shared lineage from independent evidence |
| Validation | Kd, IC50, EC50, competition or functional assays, specificity, stability, synthesis outcome | Connects selection fitness to independently measured properties |
Sequence data should distinguish nucleotide and peptide identity because synonymous templates can behave differently during transcription, translation or amplification. Modified libraries also need an explicit chemical record: NCAA identity and stereochemistry, backbone modification, cyclization anchors, ring topology and, where relevant, codon assignment and incorporation method.
Round-level data should retain R0 or another measurement of the input library whenever possible. R1, R2, R3 and later rounds are not interchangeable samples; together they show whether a sequence rises reproducibly, remains flat, drops out or appears only after a bottleneck. Selection-condition data make those changes interpretable. Enrichment against the intended target and enrichment against beads, tags, homologs or other counter-targets do not describe the same phenotype.
Finally, orthogonal validation data are especially valuable. A sequencing record reports survival through the display workflow. A measured Kd, functional response, serum-stability result or synthesis outcome records another property under another assay. Keeping those endpoints distinct prevents a model from presenting a selection signal as if it were a binding constant or a developability result.
Why Final-Round Read Count Is Not a Binding Constant
A final read count is the cumulative result of many filters:
starting abundance → transcription → translation → fusion formation → chemical maturation → target interaction → washing → recovery → amplification → sequencing
A peptide can bind well but begin at very low abundance, translate inefficiently or cyclize incompletely. Another sequence can bind only modestly yet start high, amplify efficiently or adhere nonspecifically to the matrix. Both routes alter the final count.
Therefore:
- read count is not affinity;
- enrichment is not Kd;
- the top sequence is not automatically the best peptide.
Models trained without this distinction learn a surrogate for survival through the complete protocol. That surrogate can still be useful for candidate ranking, but it must be named and evaluated as selection fitness—not silently reinterpreted as exact affinity, potency, pharmacokinetics, oral bioavailability, toxicity or human efficacy.
For a stage-by-stage explanation of these experimental filters, see From Random Peptide Libraries to Hits: What Happens in One Round of mRNA Display Selection?.
Enrichment Trajectories Carry More Information Than Raw Abundance
Relative change across rounds is usually more informative than abundance in a single pool. A rare R0 sequence that rises consistently through R1, R2 and R3 tells a different story from an initially common sequence whose normalized abundance remains flat.
A practical analysis commonly uses log enrichment or another stabilized relative-change measure. Before modeling, it should address sequencing-depth normalization, pseudocounts for zero or near-zero observations, uncertainty in low-count sequences, dropout and consistency across replicates. These steps do not make enrichment equivalent to affinity; they make the observed selection trajectory less sensitive to avoidable numerical artifacts.
| Sequence family | R0 | R1 | R2 | R3 | Counter-selection | Interpretation |
|---|---|---|---|---|---|---|
| Family A | low | ↑ | ↑↑ | ↑↑↑ | low | target-enriched |
| Family B | medium | ↑ | ↑↑ | ↑↑ | ↑↑ | possible nonspecific or off-target enrichment |
| Family C | high | stable | stable | stable | low | abundant, but not strongly enriched |
Illustrative example only; the symbols are not company experimental results.
Depletion and Counter-Selection Are Training Signals Too
Positive selection alone cannot distinguish target preference from generic stickiness. Negative and counter-selection experiments can reveal matrix binding, tag binding, off-target or homolog binding and nonspecific hydrophobic retention.
A peptide enriched on the target but also enriched on the counter-target is not equivalent to a selective binder. A useful model can compare positive and counter-selection trajectories, rank target-versus-counter-target preference, and flag families whose apparent performance is shared across controls. This turns the computational task from “find anything that survives” into “prioritize molecules whose behavior matches the intended selectivity profile.”
Experimental negatives are also more informative than arbitrary random sequences. Potential lower-fitness examples include depleted library members, non-enriched members with adequate starting counts, counter-selected sequences, experimentally inactive variants and family mutations that lose activity. Yet a non-enriched sequence is not automatically a true non-binder: sampling limits, low input representation and technical loss create uncertain labels. Training data should retain that uncertainty rather than forcing every sequence into a definitive binary class.
Ranking Is Often Better Than Binary Classification
mRNA display usually provides relative evidence. It is consequently well suited to ranking and fitness prediction, for example:
- predicting relative enrichment or the probability of continued enrichment;
- ranking candidates within a sequence family;
- ranking target-versus-counter-target preference;
- prioritizing variants for resynthesis and independent assays;
- selecting informative members for the next focused library.
These objectives match what the experiment observes more closely than a universal binder/non-binder label. They also allow teams to incorporate multiple goals later—for example selection fitness, specificity, stability and synthesizability—without pretending that one score directly measures them all.
Sequence Families Must Shape Both Analysis and Evaluation
mRNA display campaigns often produce related sequence families rather than isolated winners. Alignment within a family can reveal highly conserved residues, semi-conserved positions, tolerant positions, residue-class preferences and possible lineage-like expansion across rounds. Those patterns support constrained generation, site-specific mutation, NCAA substitution and focused-library design.
They also create a serious evaluation risk. If closely related family members are randomly divided between training and test sets, a model can appear accurate by recognizing a family it has already seen. That is data leakage, not evidence that the model can generalize to a new chemotype.
Data splitting should therefore consider sequence identity and cluster, as well as experimental round, target and campaign. Useful tests include family-held-out, campaign-held-out and, where the scientific question makes sense, target-held-out evaluation. Evaluation should test generalization, not memorization of a sequence family.
Traditional motif analysis asks which residues occur frequently at each position. AI models may additionally capture context dependence, residue interactions and non-additive or epistatic effects. These patterns can improve prediction and design, but they do not by themselves establish a causal molecular mechanism.
Chemistry-Aware Representation of NCAAs and Cyclic Peptides
When a library contains noncanonical amino acids, encoding every unusual residue as X1, X2 or one undifferentiated unknown token removes the chemistry that the experiment was designed to explore. At minimum, records should preserve NCAA identity, stereochemistry, N-methylation status, side-chain chemistry, charge, hydrophobicity, steric class, cyclization role, codon assignment and relevant incorporation method.
A chemistry-aware model may combine residue tokens with physicochemical or structural descriptors, chemical fingerprints where appropriate, backbone-modification flags and stereochemistry. It should distinguish roles such as D-amino acids, N-methyl amino acids, Aib-like constrained residues, aromatic NCAAs and cyclization handles. For the experimental basis of these libraries, see Why mRNA Display Is Especially Suited for Noncanonical Amino Acid and Cyclic Peptide Discovery and How Noncanonical Amino Acids Improve Peptide Design Beyond Stability.
Cyclic peptides require another layer. Two molecules with the same linear residue string can be different compounds if their anchor positions, closure chemistry, linker or ring topology differ. Cyclization should therefore be represented as molecular connectivity, not as a free-text label appended to a linear sequence. A minimal record includes anchor positions, cyclization chemistry, ring topology and linker identity when present.
Structure data can add target conformation, peptide conformation, predicted contacts, interface scores or pose consistency. Structure-conditioned design may help preserve hotspots or propose compatible mutations. However, a predicted peptide structure or docking pose remains a hypothesis when experimental structural evidence is unavailable, and its uncertainty should flow into downstream decisions.
Three Practical AI Tasks
1. Candidate ranking
The inputs are sequence and chemistry, the round-by-round enrichment trajectory, counter-selection data and replicate context. The output is a prioritized, auditable list for resynthesis. This is often the most direct and lowest-risk use because it narrows an existing experimental result instead of claiming to create activity from no evidence.
2. Local optimization
Starting from a validated or reproducibly enriched family, a model can prioritize nearby variants intended to improve affinity, selectivity, stability or synthesis behavior; introduce an NCAA; or compare cyclization options. Conserved positions can be protected while tolerant positions are diversified. This keeps the proposal near experimentally supported sequence space. Related downstream principles are discussed in How AI Can Optimize an Existing Peptide and Why AI-Designed Peptides Still Need Synthesizability Screening.
3. Generative design
With sufficiently broad, well-curated target–peptide, selection and validation data, generative models can propose candidates under explicit chemical constraints. One campaign is rarely enough to demonstrate a broadly capable generator. Generated molecules should be filtered for library compatibility, chemistry, topology, diversity and synthesis feasibility, then tested prospectively.
For platforms with limited data, candidate ranking, family analysis, mutation prioritization and focused-library design generally offer a more realistic first return than a large generative model.
AI Can Design the Next Library, Not Only the Next Peptide
Once multiple NCAAs and variable positions are allowed, the combinatorial space expands too quickly to enumerate. AI is useful not because it can generate more strings, but because it can reduce the search space while preserving informative diversity and chemical feasibility.
A focused library might retain a conserved binding motif, vary only tolerant positions, and allocate selected natural amino acids, D-amino acids, N-methyl residues or other permitted NCAAs according to family evidence and translation constraints. The output is not a single “best” molecule. It is a deliberately structured next experiment.
Twenty to fifty well-chosen, synthesized and tested candidates may teach the next model more than 100,000 unvalidated virtual sequences. Experimental feedback is scarce, target-specific information; virtual candidates are hypotheses. The design objective should include information gain, not only predicted score.
The Design–Build–Select/Test–Learn Loop
A closed-loop workflow can be organized as:
raw mRNA display data → QC and normalization → sequence and chemistry representation → round-by-round enrichment → positive/counter-selection comparison → family clustering → model training and ranking → candidate or focused-library design → synthesis, display or independent assay → experimental feedback → model update

Figure 1. mRNA display data and AI peptide design form an experimental feedback loop. The model prioritizes testable hypotheses; selection and independent assays produce different but complementary feedback.
This is a Design → Build → Select/Test → Learn → Redesign cycle. The distinction between selection and an independent assay matters. Re-entering a focused library into mRNA display produces another selection-fitness trajectory; resynthesizing discrete peptides and measuring binding or function produces orthogonal labels. Both can update a model, but they should not be merged without endpoint metadata.
Data Quality and Provenance Come Before Model Complexity
A reusable data set needs stable sequence identifiers, normalized molecular representations, modification and topology records, round and target IDs, selection type, counts and sequencing depth, replicate IDs, library chemistry, assay results and provenance.
Every training example should be traceable to its campaign, round, target, positive or negative selection, library, chemistry, assay and processing version. It should also state whether a value is experimental, derived or simulated. Badly structured experimental data cannot be rescued by a more complex model.
Simulated data can help develop a pipeline, test schemas, prototype model architectures and validate workflow mechanics. It cannot demonstrate real discovery performance in the chemistry and bias structure of an mRNA display campaign. Synthetic and experimental records must remain distinguishable, and claims about discovery capability must ultimately rest on real selection and assay data.
Evaluation Must End in Prospective Validation
Historical test-set performance is useful but insufficient. Random splits can inflate results through family overlap or campaign-specific shortcuts. Family-held-out and campaign-held-out tests are stronger, and target-held-out evaluation can be appropriate when the intended use is transfer to new targets.
The strongest evidence is prospective: the model prioritizes candidates before their outcomes are known; those candidates are synthesized, selected or assayed; and their hit rate, enrichment, selectivity or another predefined endpoint is compared with a relevant baseline. This asks the operational question that matters—did the model improve the next experimental decision?
From a Selection Signal to a Developable Peptide
High enrichment does not guarantee easy synthesis, solubility, proteolytic or serum stability, selectivity, low toxicity or permeability. After selection-based prioritization, development requires additional evidence and often a multi-objective design step incorporating physicochemical properties, structure, synthesis feasibility and independent assay results.
AI should not claim exact Kd, IC50, PK, oral bioavailability, toxicity or human efficacy from sequencing trajectories alone. Its most direct role is selection-fitness analysis, candidate prioritization and experimental design. Custom peptide and cyclic peptide synthesis can then convert encoded or designed candidates into physical samples for identity, purity, binding and functional testing.
Conclusion
The real integration of mRNA display and AI is not the act of handing final-round sequencing results to a model. It is the conversion of the entire selection process into structured, traceable data with experimental context.
When enrichment, depletion, sequence families, positional tolerance, noncanonical amino acids, cyclization topology and independent validation can be represented together, AI can move from analyzing a completed selection toward designing the next experiment. That transition should begin with careful candidate ranking and focused-library design, and it should be judged by prospective laboratory results.
References
- Roberts RW, Szostak JW. RNA-peptide fusions for the in vitro selection of peptides and proteins. PNAS. 1997;94:12297–12302. doi:10.1073/pnas.94.23.12297
- Blanco C, Verbanic S, Seelig B, Chen IA. High throughput sequencing of in vitro selections of mRNA-displayed peptides: data analysis and applications. Physical Chemistry Chemical Physics. 2020;22:6492–6506. doi:10.1039/C9CP05912A
- Kamalinia G, et al. Directing evolution of novel ligands by mRNA display. Chemical Society Reviews. 2021;50:9055–9103. doi:10.1039/D1CS00160D
- Josephson K, Ricardo A, Szostak JW. mRNA display: from basic principles to macrocycle drug discovery. Drug Discovery Today. 2014;19:388–399. doi:10.1016/j.drudis.2013.10.011
- Li S, Millward S, Roberts RW. In Vitro Selection of mRNA Display Libraries Containing an Unnatural Amino Acid. Journal of the American Chemical Society. 2002;124:9972–9973. doi:10.1021/ja026789q
- Huang Y, Wiedmann MM, Suga H. RNA display methods for the discovery of bioactive macrocycles. Chemical Reviews. 2019;119:10360–10391. doi:10.1021/acs.chemrev.8b00430
- Vinogradov AA, Chang JS, Onaka H, Goto Y, Suga H. Accurate Models of Substrate Preferences of Post-Translational Modification Enzymes from a Combination of mRNA Display and Deep Learning. ACS Central Science. 2022;8:814–824. doi:10.1021/acscentsci.2c00223
- Smith TP, Bhushan B, Granata D, et al. Identification and engineering of potent cyclic peptides with selective or promiscuous binding through biochemical profiling and bioinformatic data analysis. RSC Chemical Biology. 2024;5:12–18. doi:10.1039/D3CB00168G
Want to use mRNA display results for the next candidate design round?
If your project has generated mRNA display, NGS, or candidate peptide data, we can analyze sequence families, enrichment trajectories, NCAA chemistry, cyclization, and downstream validation to support the next design round and custom synthesis.