From Random Peptide Libraries to Hits: What Happens in One Round of mRNA Display Selection?
One round of mRNA display is a population process shaped by library design, translation, fusion formation, target selection, counter-selection, washing, recovery, amplification, and NGS.
From Random Peptide Libraries to Hits: What Happens in One Round of mRNA Display Selection?
A common shorthand describes display screening as “finding the strongest binders in a large library.” An mRNA display selection is more accurately understood as a population process. A complex molecular population passes through a series of transformations and selection pressures, and every stage can change which sequences remain observable.
Final abundance is influenced by initial representation, transcription, translation, fusion formation, noncanonical amino acid (NCAA) incorporation, chemical maturation, target binding, nonspecific retention, washing, recovery, amplification, and sequencing depth. A sequence that becomes abundant may be a useful binder, but abundance alone does not prove high affinity.
This article follows a representative selection round at a conceptual level. Exact ordering varies among platforms: reverse transcription may occur before or after target selection, and chemical maturation may be unnecessary for a canonical linear library. The central principle remains the same: the recovered sequence distribution is a composite output of selection and molecular processing.
The Workflow at a Glance
| Stage | What happens | What changes in the population |
|---|---|---|
| Library design | Variable positions, codons, fixed motifs, NCAA assignments, and cyclization anchors are defined | The accessible sequence and chemical space is bounded before the experiment starts |
| Transcription | DNA templates are converted into RNA | Yield, truncation, and RNA integrity can alter representation |
| Translation | RNA templates produce peptides in a cell-free system | Codon context, stalling, tRNA supply, and monomer compatibility affect peptide output |
| Fusion formation | The peptide becomes covalently linked to its encoding mRNA, commonly through a puromycin-based architecture | Unequal fusion efficiency determines which genotype–phenotype pairs enter selection |
| Chemical maturation | Optional deprotection, cyclization, oxidation, or chemoselective conversion creates the intended displayed species | Incomplete or sequence-dependent conversion creates product heterogeneity |
| Target selection | The library contacts a presented target | The target state and selection conditions define which interactions are favored |
| Washing and counter-selection | Weak, fast-dissociating, matrix-binding, tag-binding, or off-target species are differentially removed | Stringency and control design shape signal-to-background and specificity |
| Recovery | Retained genotype-linked molecules are collected | Elution or recovery efficiency can favor some retained species over others |
| RT–PCR and amplification | Recovered sequence information is converted and amplified | Template structure, GC content, stochastic bottlenecks, and PCR efficiency can distort abundance |
| NGS and analysis | Sequence counts, families, enrichment, and depletion are measured | Read depth and analysis choices determine which trajectories can be resolved |

Figure 1. A representative mRNA display selection cycle. The arrows describe the experimental flow, but the output of every stage is also a new population shaped by technical and biochemical filters.
1. Selection Begins with Library Design
Selection pressure starts before the library contacts a target. Random-region length, codon design, fixed positions, scaffold choice, cyclization anchors, NCAA codon allocation, and sequence constraints determine what can exist in the experiment.
Theoretical diversity is not physical diversity. A degenerate design may describe far more possible sequences than the number of molecules that can be synthesized and handled. Even when each residue is intended to occur equally, oligonucleotide synthesis and template construction can create unequal starting representation. A sequence absent from the physical input cannot be recovered by later selection.
Good analysis therefore benefits from a round-zero measurement or another characterization of the input pool. Enrichment should be interpreted relative to where a sequence started, not only by where it finished.
For the principles behind NCAA-rich and cyclic libraries, see Why mRNA Display Is Especially Suited for Noncanonical Amino Acid and Cyclic Peptide Discovery. A concise overview of genotype–phenotype linkage is available in mRNA Display: Discovering High-Affinity Peptides from Ultra-Large Libraries.
2. Transcription Changes Representation Before Selection
DNA templates are transcribed into an RNA library. Sequence-dependent yield, incomplete products, and differences in RNA integrity can distort the population. These effects are not target-binding events, but they affect which templates reach translation.
This creates a one-way constraint: if a sequence is lost or severely underrepresented during transcription and purification, later selection cannot reveal its theoretical binding potential. Quality control at the nucleic-acid stage is therefore part of interpreting the final peptide population, even though it is not itself an affinity-selection step.
3. The Translated Library May Differ from the Designed Library
Translation efficiency is not identical across all templates. Codon context, difficult sequence features, ribosome stalling, charged-tRNA availability, and peptide properties can affect productive synthesis. In an NCAA library, a codon assignment in the design does not guarantee the same incorporation frequency in every sequence context.
Aminoacylation efficiency, tRNA compatibility, delivery to the ribosome, and ribosomal acceptance all contribute. Repeated or adjacent incorporation can behave differently from a single NCAA placement. The relevant experimental population is therefore the translated peptide library, not merely the designed DNA library.
4. Fusion Formation Defines Which Molecules Are Traceable
mRNA display depends on genotype–phenotype linkage. A common architecture uses a puromycin-containing linker so that the nascent peptide becomes covalently connected to its encoding mRNA during translation. This linkage makes a selected phenotype recoverable through its genotype.
Fusion formation is also a filter. If templates differ in their probability of producing stable mRNA–peptide fusions, some sequences may be systematically underrepresented before target incubation. Measuring or controlling overall fusion performance helps distinguish a weak campaign from a poor fusion-preparation step; it does not, however, guarantee equal fusion efficiency for every member.
5. Optional Chemical Maturation Creates the Screened Molecule
Translation is not always the final library-construction step. NCAA-containing or cyclic peptide libraries may require deprotection, spontaneous or induced cyclization, oxidation, chemoselective ligation, or another compatible post-translational conversion.
Selection should act as far as practical on the intended mature chemical species. If ring closure is incomplete or sequence-dependent, the pool may contain linear precursor, desired macrocycle, and side products that experience different binding and recovery behavior. Apparent sequence enrichment may then combine biological recognition with chemical-conversion efficiency.
This is why the interpretation of NCAA and cyclic libraries requires information beyond sequence identity: codon-reassignment fidelity, incorporation efficiency, cyclization conversion, and mature-product heterogeneity may all matter.
6. Target Incubation Defines the Selected Phenotype
The mature fusion library is incubated with an immobilized or otherwise presented target. Target concentration, presentation, accessibility, conformation, and the incubation regime collectively determine which binders can be retained.
Higher target concentration may preserve weaker interactions; lower concentration may increase pressure but can also reduce recoverable signal. Likewise, changing contact time may favor different kinetic behaviors. There is no universal rule that lower target concentration is always better. The conditions should reflect the intended selection objective and the ability to distinguish signal from background.
Target quality matters
A selection cannot compensate for an inappropriate target preparation. Folding, oligomeric state, post-translational modifications, active conformation, immobilization orientation, and tag exposure all require consideration. A surface-presented target may unintentionally select binders to the matrix, linker, tag, exposed hydrophobic patches, or a non-native conformation.
Target controls are therefore experimental information, not administrative details. They define what “specific binding” can mean in that campaign.
7. Positive Selection and Counter-Selection Serve Different Purposes
Positive selection retains library members that survive interaction with the intended target. Used alone, it can also enrich bead binders, resin binders, tag binders, hydrophobic sticky peptides, nucleic-acid-interacting artifacts, or molecules that recognize unrelated proteins in the matrix.
Negative or counter-selection removes specified unwanted populations. A library may first contact a blank matrix to deplete matrix-binding sequences, or contact a homolog or off-target to reduce cross-reactive families before positive selection against the intended target.
Specificity is therefore partly designed into the selection strategy. A positive selection against target A combined with depletion against target B may favor A-selective behavior, but it cannot by itself establish specificity. Independent assays against A, B, and relevant controls are still required.
8. Washing Is a Tunable Selection Pressure
Washing determines which bound species survive. It can reduce nonspecific background and remove weak or fast-dissociating interactions, but wash stringency is not a simple “more is better” variable.
Excessive stringency can discard useful binders, especially early in a campaign or when target density and recovery are limiting. Insufficient stringency can allow sticky peptides and background interactions to dominate amplification. The experimental problem is signal-to-background optimization under the intended binding model.
9. Recovery Adds Another Filter
Retained genotype-linked molecules must be recovered. Recovery efficiency can vary: weak binders may already have been lost during washing, very persistent interactions may be difficult to release, and some molecules may remain through nonspecific mechanisms.
The recovered pool is therefore not a pure measurement of equilibrium affinity. It is the product of target association, dissociation during washing, nonspecific retention, and the recovery method.
10. Amplification Preserves Information—and Can Distort It
Recovered genetic material is reverse-transcribed and amplified to produce material for the next round and/or sequencing. PCR is essential because the selected population is small, but different templates can amplify with different efficiencies because of GC content, secondary structure, primer interactions, and other sequence properties.
Stochastic bottlenecks are also important when few molecules are recovered. A sequence can appear to expand because of early sampling effects, while a true binder can be lost by chance. Repeated amplification across rounds can compound these differences. Final read count must therefore not be interpreted directly as affinity or dissociation constant.
11. NGS Measures a Composite Observable
Next-generation sequencing (NGS) can measure sequence abundance, enrichment, depletion, family structure, motif emergence, and positional preferences. Its depth makes it possible to observe trajectories that would be invisible by low-throughput sequencing.
Yet an NGS count is a composite observable. It reflects input representation, molecular processing, selection, recovery, amplification, sampling depth, and data processing. The highest-count sequence is not automatically the best binder.
Useful interpretation integrates several signals:
- abundance in the starting and intermediate pools;
- fold enrichment across adjacent rounds;
- consistency across replicates;
- membership in an independently enriched sequence family;
- behavior in counter-selection or off-target controls;
- sequencing depth and uncertainty;
- independent biochemical validation.
One Round Is a Snapshot, Not a Complete Campaign
Early-round pools often contain high background and rare functional sequences mixed with technical noise. Iterative rounds use the recovered and amplified pool as the next input while selection pressure may be adjusted. Useful families can become more visible over time, but additional rounds are not automatically beneficial: excessive convergence can remove diversity, magnify amplification artifacts, or favor the easiest-to-propagate sequences.
A trajectory such as low abundance in round 1 followed by consistent increases in rounds 2, 3, and 4 is often more informative than a sequence that appears at high abundance only in the final round. Even then, interpretation should consider starting abundance, enrichment ratio, replicate consistency, family context, negative-selection behavior, and sequencing coverage.
Sequence Families Are Often More Informative Than a Top-10 List
Selected pools commonly contain related sequence families rather than only isolated winners. Family members may share a binding motif, conserved charge or hydrophobic pattern, ring topology, or NCAA position.
Highly conserved positions may contribute to target contact, structural organization, or cyclization geometry. Variable positions can indicate tolerance and provide practical sites for structure–activity relationship (SAR) exploration. A family-level view can therefore distinguish a reproducible chemical pattern from a single high-count outlier.
Positive enrichment and negative depletion should also be analyzed together. A family enriched against the desired target but equally enriched against a counter-target is not strong evidence of specificity.
Replicates Improve Confidence
Selection and sequencing can contain handling variation, amplification bias, and stochastic bottlenecks. Where feasible, replicate selections or strategically repeated stages help reveal which trajectories are reproducible. Replication does not remove bias automatically, but it helps separate stable biological signals from one-off population events.
Sources of Bias Across the Workflow
| Stage | Potential bias | Interpretation risk |
|---|---|---|
| Library synthesis | Uneven starting representation | Rare or absent designs cannot be fairly compared |
| Transcription | Sequence-dependent yield, truncation, RNA integrity | Nucleic-acid loss can resemble negative selection |
| Translation | Codon context, stalling, sequence-dependent yield | Designed abundance differs from peptide abundance |
| NCAA incorporation | Variable charging and incorporation efficiency | Codon presence does not prove mature NCAA incorporation |
| Fusion formation | Unequal fusion efficiency | Low-fusion members are systematically undercounted |
| Cyclization or maturation | Incomplete or sequence-dependent conversion | Reads may represent heterogeneous chemical species |
| Target binding | Target state, presentation, concentration, accessibility | Enrichment may favor an unintended target state |
| Washing | Stringency and dissociation-time bias | Useful binders may be lost or background retained |
| Recovery | Unequal release or elution | Retention does not guarantee efficient recovery |
| RT–PCR | Template structure, GC content, stochastic bottlenecks | Amplification advantage can mimic enrichment |
| Sequencing | Sampling depth, errors, preprocessing choices | Low-frequency trajectories may be missed or misestimated |
A Selection Hit Is Not Yet a Validated Free Peptide
A selected sequence should be independently resynthesized, its mass and purity confirmed, and its binding tested outside the display construct. Competition, functional, specificity, stability, and project-appropriate developability assays may then be required.
Resynthesis breaks the genotype–phenotype system and tests the peptide as an independent chemical entity. This matters because the free peptide can behave differently from an mRNA-linked species, and because the chemical product obtained by synthesis may expose ambiguities in cyclization or NCAA incorporation.
A practical path is:
enriched sequence → independent synthesis → binding validation → family analysis → SAR → NCAA and cyclization refinement → stability, solubility, and synthesizability assessment → next-generation candidate
mRNA display is a discovery engine, not the complete drug-development process. Downstream design considerations are discussed in How AI Can Optimize an Existing Peptide, Why AI-Designed Peptides Still Need Synthesizability Screening, and the Custom Peptide Synthesis Guide.
Why Selection Trajectories Are Useful for AI
The most useful computational dataset is not merely a list of final winners. It can include the initial sequence, round-by-round abundance, positive enrichment, negative depletion, family membership, positional conservation, library chemistry, and selection conditions.
Such longitudinal and structured data may support candidate ranking, sequence–fitness modeling, motif discovery, experimental prioritization, and generative peptide modeling. These applications still require careful normalization, uncertainty estimates, chemically aware representations, and experimentally validated labels. A later article will discuss how mRNA display data can be integrated with AI-assisted peptide design.
Conclusion
The value of mRNA display is not simply that a few abundant sequences emerge from a very large library. The scientifically useful question is how peptide families change under continuous selection pressure, and which sequence and chemical features are associated with target binding, selectivity, and chemical executability.
From library composition, translation, fusion formation, and cyclization to target selection, counter-selection, PCR, and NGS, every stage shapes the observed data. Understanding these filters is what allows an enriched sequence to become a candidate worth resynthesizing and validating.
References
- Roberts RW, Szostak JW. RNA-peptide fusions for the in vitro selection of peptides and proteins. PNAS. 1997;94:12297–12302. doi:10.1073/pnas.94.23.12297
- Wilson DS, Keefe AD, Szostak JW. The use of mRNA display to select high-affinity protein-binding peptides. PNAS. 2001;98:3750–3755. doi:10.1073/pnas.061028198
- Kamalinia G, et al. Directing evolution of novel ligands by mRNA display. Chemical Society Reviews. 2021;50:9055–9103. doi:10.1039/D1CS00160D
- Newton MS, Cabezas-Perusse Y, Tong CL, Seelig B. In Vitro Selection of Peptides and Proteins—Advantages of mRNA Display. ACS Synthetic Biology. 2020;9:181–190. doi:10.1021/acssynbio.9b00419
- Blanco C, Verbanic S, Seelig B, Chen IA. High throughput sequencing of in vitro selections of mRNA-displayed peptides: data analysis and applications. Physical Chemistry Chemical Physics. 2020;22:6492–6506. doi:10.1039/C9CP05912A
- Li S, Millward S, Roberts RW. In Vitro Selection of mRNA Display Libraries Containing an Unnatural Amino Acid. Journal of the American Chemical Society. 2002;124:9972–9973. doi:10.1021/ja026789q
- Peacock H, Suga H. Discovery of De Novo Macrocyclic Peptides by Messenger RNA Display. Trends in Pharmacological Sciences. 2021;42:385–397. doi:10.1016/j.tips.2021.02.004
- Nishikawa S, et al. De Novo Single-Stranded RNA-Binding Peptides Discovered by Codon-Restricted mRNA Display. Biomacromolecules. 2024;25:355–365. doi:10.1021/acs.biomac.3c01024
Have a target and want to begin peptide discovery or candidate optimization?
If your project involves a protein target, mRNA display, cyclic peptides, noncanonical amino acids, or an existing peptide hit, we can support candidate design, chemical feasibility assessment, custom synthesis, purification, and analytical quality control.