Why AI-Designed Peptides Still Need Synthesizability Screening
AI peptide design can generate candidates quickly, but biological scores do not guarantee practical synthesis or purification. This article explains how peptide synthesizability, SPPS, noncanonical amino acids, cyclic peptides, purification difficulty, and practical peptide synthesis should shape candidate selection.
Why AI-Designed Peptides Still Need Synthesizability Screening
Generative models and sequence optimization models can propose dozens, hundreds, or even thousands of candidate peptides in a very short time. The increase in the number of candidates expands the search space, but it does not automatically improve the experimental success rate. Before placing an order for synthesis, the project team still needs to answer a set of very specific questions: Is the sequence likely to aggregate on the resin? Does it contain consecutive difficult coupling sites? Are deletion peptides, truncated peptides, or specific side reactions likely to increase significantly? Are suitable protected building blocks available for the NCAAs, and can they be procured? Is the envisioned cyclization route compatible with the protecting group strategy? Can the target peak in the crude product be separated from closely related impurities?
These questions determine whether a computational candidate can enter real experiments at a reasonable cost. A sequence may perform well in activity, affinity, or stability models, yet become a high-cost, low-output, or even temporarily infeasible project because of synthesis and separation risks. For peptides, synthesizability is part of the design objective, not a supplementary check after design is complete.
Why Biological Scores Do Not Equal Synthetic Feasibility
The objective functions of most sequence models focus on metrics such as activity, affinity, toxicity, solubility, stability, permeability, or sequence likelihood. These objectives are important for candidate discovery, but they do not inherently include coupling efficiency, on-resin aggregation, side reactions, crude purity, cyclization efficiency, and purification difficulty in solid-phase peptide synthesis (SPPS). The reason is not simply that “AI does not understand chemistry”; rather, what a model learns depends on what is recorded in the training data and optimization objectives.
For example, a model may favor a hydrophobic interface because it facilitates interactions with the target protein; it may also recommend N-methylation or bulky residues to stabilize a particular binding conformation. However, these modifications may also reduce chain segment accessibility in the resin swelling environment, increase steric hindrance during coupling, or make the chromatographic behavior of the target product very similar to that of deletion peptides. If the training data contain only biological activity labels and no actual synthesis batches, crude-product chromatograms, or separation results, the model does not have sufficient information to independently assess these consequences.
Therefore, after AI-assisted peptide design, chemical-level candidate review is still required. Synthesizability screening does not negate the biological value provided by the model; instead, it further converts “worth testing” into “worth doing first, feasible to make, and how to make it.”
On-Resin Aggregation: Risks Beyond Solution Properties
Consecutive hydrophobic residues such as Leu, Ile, Val, Phe, and Trp, or long segments that readily form stable secondary structures, may cause the growing peptide chain to aggregate on the resin. Aggregation reduces the accessibility of reaction sites to amino acid monomers and coupling reagents, making otherwise routine condensations slow or incomplete. After unreacted chains continue into the next cycle, sequences missing one or more residues are formed; the problem gradually accumulates in long chains and ultimately manifests as a complex crude product and a reduced proportion of the main peak.
High hydrophobicity does not necessarily mean that a peptide cannot be synthesized, and a single hydrophobic residue should not be mechanically classified as a failure. Risk depends on residue distribution, sequence length, charge, resin and linker, protecting groups, solvent, temperature, and the synthesis strategy used. A reasonable screening system is better suited to identifying consecutive hydrophobic segments, possible self-association regions, and risk combinations, and then using them as the basis for candidate ranking and route design, rather than applying an absolute threshold.
Steric Hindrance and Difficult Couplings Often Come from Valuable Designs
β-branched residues, N-methyl amino acids, α,α-disubstituted amino acids, and some bulky NCAAs can all increase local coupling difficulty. When highly hindered residues appear consecutively, the nucleophilicity of the amino group of the preceding residue, the ability of the next activated monomer to approach the reaction site, and the local conformation of the peptide chain may all be affected simultaneously. The substitutions recommended by a model may indeed contribute to conformational preorganization, protease stability, or the binding interface, but they may also require longer reaction times, repeated couplings, different activation systems, or fragment condensation strategies.
N-methyl amino acids are a typical example. N-methylation can alter the number of hydrogen-bond donors and the local conformation, but the steric hindrance and lower reactivity of the N-methylamino group also affect subsequent bond formation. During screening, both its design value and its synthesis cost must be retained: candidates should not be deleted as soon as an N-methyl residue is encountered, nor should it be treated as an ordinary character with no process impact.
Known Side Reactions Need to Be Included in Candidate Ranking
Risks in SPPS are often jointly determined by “residue + neighboring environment + operating conditions.” Asp-related sequences may undergo side reactions via the aspartimide pathway under basic deprotection conditions; certain N-terminal dipeptides carry a risk of diketopiperazine formation after deprotection; Cys may undergo oxidation or form mismatches in multi-disulfide systems; Met and Trp require attention to oxidation. Some activation and coupling environments may also increase the risk of racemization. Deamidation issues for Asn and Gln are more related to subsequent handling, storage conditions, and the sequence microenvironment, and they likewise affect the stability assessment of the final sample.
These risks do not mean that the relevant residues cannot be used. Their significance lies in reminding the design team that, among two candidates with similar biological scores, if one has multiple overlapping difficult motifs while the other can yield an analyzable sample through a more straightforward route, the latter is usually more suitable for entering the first round of experiments. Synthesizability screening should identify the sources of risk so that chemists can select protecting groups, capping, oxidation control, or alternative routes, rather than outputting only an opaque “pass/fail” result.
Noncanonical Amino Acids Are Not Ordinary Characters in an Expanded Alphabet
Noncanonical amino acids (NCAAs) can expand chemical space and introduce new side-chain interactions, conformational constraints, and metabolic stability, making them important tools in modern peptide design. From a synthesis perspective, however, an NCAA cannot be treated merely as a new symbol outside the natural amino acid alphabet. In actual projects, at least the following levels need to be checked:
| Evaluation Level | Questions to Answer |
|---|---|
| Material | Is the protected building block with the corresponding configuration and purity commercially available, and are the lead time and cost suitable for the project? |
| Protecting group | Is side-chain protection compatible with the Fmoc/Boc route, cleavage conditions, and other sensitive groups? |
| Reactivity | Are monomer activation, coupling, and subsequent residue extension affected by electronic effects or steric hindrance? |
| Post-processing | Is deprotection complete, and is it prone to oxidation, racemization, or formation of byproducts that are difficult to separate? |
| Downstream steps | Does this residue affect cyclization, disulfide bond formation, purification, and mass spectrometric confirmation? |
D-amino acids are often used to modulate conformation or protease stability; Aib has pronounced conformational preferences and steric hindrance; Nal increases the aromatic hydrophobic surface; Orn and Dab can provide side-chain amino groups of different lengths and participate in modification or lactam cyclization. Each type of structure has design value, as well as different requirements for protected building blocks, coupling, and purification. A more complete design process simultaneously evaluates “whether this residue contributes to the target properties” and “whether this specific building block can be reliably used in the current route.” For more background, see the role of NCAAs in drug development.
The Cyclization Method Itself Is Part of Synthesizability
When AI designs cyclic peptides, it cannot determine whether two sites can be “connected” solely from a structural diagram. Disulfide cyclization, head-to-tail cyclization, lactam cyclization, side-chain-to-side-chain cyclization, and side-chain-to-C-terminus cyclization each have different requirements for precursor design, orthogonal protection, reaction concentration, and purification. Ring size, residue geometry, and conformational preorganization influence intramolecular reactions; if the precursor cannot approach itself effectively, intermolecular coupling, oligomerization, epimerization, or other side reactions may increase.
Three-dimensional structures and conformational sampling can help compare candidates, but actual cyclization yield cannot be accurately predicted solely from a single predicted conformation. A reasonable approach is to incorporate cyclization type, site geometry, protecting group compatibility, precursor solubility, and potential competing reactions into the evaluation together, while retaining alternative ring-closure routes. The cyclic peptide design guide discusses the drug design value brought by cyclization; in project execution, these values need to be assessed together with practically implementable synthetic routes.
Being Able to Synthesize It Does Not Mean Being Able to Purify It
What commercial projects truly care about is not only whether the target bond is formed, but also whether the target compound can be separated from the crude product to meet quality requirements. Deletion peptides and the main product may differ by only one residue and have very similar retention behavior; extremely hydrophobic sequences may be difficult to dissolve and strongly retained, while extremely hydrophilic sequences may lack a sufficient separation window in reversed-phase chromatography. Oxidized forms, epimers, cyclization byproducts, and multiple disulfide-bond isomers further increase peak shape and component complexity.
This is also why crude purity, purification difficulty, and expected isolated yield are closer to real-world business synthesizability than simply judging “whether the reaction occurs.” Two candidates that both meet analytical purity requirements may require completely different numbers of preparative runs, solvent systems, and process development investments. If the first round of design can avoid unnecessary closely related impurities, extreme physicochemical properties, or risks of multiple isomers, more resources can be devoted to truly informative experimental comparisons. For the connection between synthesis and quality control in routine projects, see the custom peptide synthesis guide.
From Binary Judgment to Explainable Integrated Evaluation
A more reasonable workflow is not “generation → ranking only by activity → synthesis,” but rather:
Candidate generation → Biological property screening → Chemistry and synthesizability screening → Structural evaluation → Candidate diversity selection → Synthesis and QC → Experimental feedback
Synthesizability should also not consist only of PASS or FAIL. Projects can classify candidates as “high feasibility,” “moderate feasibility,” and “challenging,” and indicate whether the main risks arise from on-resin aggregation, difficult coupling, NCAA materials, cyclization, or purification. If a challenging candidate has a unique biological hypothesis, it can still be retained, but alternative sequences, routes, and budgets should be prepared in advance; when multiple candidates have similar value, priority should be given to combinations that provide high information gain and lower chemical risk.
This approach also preserves candidate diversity. If only the sequences with the highest overall scores and high similarity to one another are selected, a single failure may not reveal whether the problem comes from the target hypothesis, a shared motif, or the synthetic route. Selecting several candidates that occupy different design trade-offs is often more suitable for early validation than pursuing a single “champion sequence.”
Peptide Design Is Fundamentally Multi-Objective Optimization
Affinity, stability, solubility, and permeability are intended to be increased, while toxicity, hemolysis, and synthesis difficulty are intended to be reduced, but these objectives do not always move in the same direction. Increasing hydrophobicity may improve certain binding interfaces or membrane interactions while also bringing aggregation, nonspecific binding, and purification problems; N-methylation may improve stability and permeability, but may also increase coupling difficulty or alter key hydrogen bonds; cyclization may stabilize the active conformation, but it increases the complexity of route control and isomer control.
Therefore, in real AI-driven peptide design, there is no universal overall score that is detached from the project context. A Pareto-based approach is better suited to expressing this type of problem: retaining a set of solutions that offer reasonable trade-offs among different objectives and are not comprehensively dominated by another candidate, and then selecting the first-round combination according to the experimental objective. Synthesizability is one of these objectives, and together with candidate diversity, budget, delivery timeline, and available experimental information, it determines prioritization.
What Can Be Evaluated, and What Still Must Be Confirmed Experimentally
Based on sequence, structural constraints, and materials information, factors that can be used for candidate prioritization include composition, hydrophobicity, charge, aggregation propensity, known difficult motifs, modification complexity, availability of NCAA protected building blocks, complexity of the cyclization route, and estimated difficulty of synthesis and purification. These metrics can help identify obvious risks, compare relative priorities, and provide input for synthetic route planning.
However, computational screening cannot reliably replace actual batch data. It cannot accurately determine isolated yield, exact crude purity, coupling efficiency at each step, or cyclization yield from sequence alone, nor can it prove actual binding affinity and biological activity. Even if historical data and models can provide probabilities or intervals, the final conclusions still need to be confirmed through synthesis, purification, mass spectrometry and chromatographic quality control, and the corresponding biological experiments. Scientifically disciplined differentiation between “risk ranking” and “outcome prediction” can prevent design confidence from being mistaken for experimental fact.
Why Design-to-Synthesis Can Shorten the Validation Cycle
When the design team and the synthesis team are completely disconnected, common outcomes include NCAAs in the sequence that cannot be procured or have excessively long lead times, unnecessary high-cost modifications, cyclization sites lacking feasible protection strategies, or the discovery only after the crude product appears that the major impurities are difficult to separate. The project then returns to the design side for modification, synthesis is scheduled again, and the design-build-test cycle is repeatedly extended.
Design-to-Synthesis emphasizes completing chemical review before candidate freeze and feeding synthesis, purification, and QC results back into the next round of design:
Design → chemistry review → synthesis feasibility → synthesis → purification and QC → experimental feedback → redesign
Apollomics is integrating sequence- and structure-based peptide design with practical peptide synthesis, purification, and QC workflows. The goal is not to promise that every computational candidate can be completed at a fixed yield, but rather to identify risks earlier, compare executable routes, and improve the efficiency with which valuable candidates enter the experimental validation stage.
Conclusion
For peptides, excellent AI design must not only answer “whether this sequence may have activity,” but must also further answer “whether this sequence can be reliably synthesized, purified, and validated at a reasonable cost.” Synthesizability is not a checkpoint after design is completed; it should be incorporated into the design objectives from the stages of candidate generation, screening, and combination selection.
Only by connecting biological prediction, chemical feasibility, preparative separation, and experimental feedback into a closed loop can the large number of rapidly generated sequences truly be translated into high-quality experimental candidates.
References
- Behrendt R, White P, Offer J. Advances in Fmoc solid-phase peptide synthesis. Journal of Peptide Science. 2016;22:4–27. doi:10.1002/psc.2836
- Coin I, Beyermann M, Bienert M. Solid-phase peptide synthesis: from standard procedures to the synthesis of difficult sequences. Nature Protocols. 2007;2:3247–3256. doi:10.1038/nprot.2007.454
- Muttenthaler M, King GF, Adams DJ, Alewood PF. Trends in peptide drug discovery. Nature Reviews Drug Discovery. 2021;20:309–325. doi:10.1038/s41573-020-00135-8