The largest number in this study is 40,626,260. The most important one may be 58.
A University of Pennsylvania team used a deep-learning system to examine more than 40 million short peptide sequences generated from venom proteins. The computer completed that search in hours and eventually produced a shortlist of 386 unusually novel candidates. Researchers could not test millions of molecules, but they could synthesize 58. Of those, 53 inhibited at least one bacterial strain in the laboratory.
That result offers a sharp example of what artificial intelligence can contribute to drug discovery. It can make an impossible search selective. It does not turn a prediction into a medicine, remove toxicity or replace the years of evidence between an active molecule and an approved antibiotic.
The distinction matters because the need is real. The latest global burden analysis estimated that bacterial antimicrobial resistance was directly responsible for 1.14 million deaths in 2021. Another 4.71 million deaths were associated with resistant infections, a broader category in which resistance was present but not necessarily the decisive cause. The Institute for Health Metrics and Evaluation explains the two counterfactual estimates and why they should not be collapsed into one figure.
Meanwhile, the World Health Organization’s 2025 review of the antibacterial pipeline describes a continuing void in development. A promising lead from venom is therefore worth attention, but its stage of development must remain visible.
The 40.6 million sequences did not come from 40.6 million vials
The team began with 16,123 venom proteins collected from four specialist databases. Together they covered proteins associated with snakes, spiders, scorpions, cone snails, sea anemones and insects.
Software then moved a sliding window along each protein sequence. Every eligible string between eight and 50 amino acids long became a possible “venom-encrypted peptide”, or VEP. This produced the headline library of 40,626,260 sequences.
That wording is important. The study did not extract 40.6 million distinct compounds from venom glands. It computationally generated short fragments encoded inside known venom proteins. Some may become biologically useful when separated from their parent protein, but the paper does not establish that every fragment exists naturally as a free molecule inside an animal.
This is less like searching 40 million bottles on a shelf than cutting millions of possible phrases from a library of longer biological sentences. The question was whether any of those phrases could act independently against bacteria.
APEX turned one enormous library into a testable shortlist
The deep-learning model, called APEX, had been trained on an in-house peptide dataset and publicly catalogued antimicrobial peptides. For a supplied amino-acid sequence, it predicts minimum inhibitory concentrations, or MICs, against 34 bacterial strains. MIC is the lowest tested concentration that prevents visible bacterial growth under specified conditions.
According to the open-access study in Nature Communications, APEX identified 7,379 venom-derived sequences with a median predicted MIC of 32 micromoles per litre or lower. The researchers then imposed a novelty filter. A candidate had to share less than 75 percent sequence similarity with known antimicrobial peptides and with other selected candidates.
Synthesis feasibility and the tendency to aggregate provided additional filters. The final 386 were not merely the highest predicted scorers. They were chosen to combine predicted potency with distance from familiar antimicrobial sequences and a practical chance of being made and handled.
That distance is valuable because resistance can travel across related drugs. Novel sequence does not guarantee a novel clinical mechanism, but repeatedly rediscovering close variations of known molecules would do less to expand the search space.
Fifty-three laboratory hits, with a denominator that matters
The team bought 58 synthesized peptides from the shortlist and tested them through broth microdilution. The panel included Acinetobacter baumannii, Klebsiella pneumoniae, two Pseudomonas aeruginosa strains, ordinary and methicillin-resistant Staphylococcus aureus, colistin-resistant Escherichia coli, and vancomycin-resistant enterococci.
Fifty-three of the 58 showed potent activity against at least one tested pathogen. The University of Pennsylvania described the computational scan as taking hours, compared with the years that a conventional search across this much sequence space would demand.
The apparent hit rate is about 91 percent, but it has a specific denominator. These were 58 promising candidates chosen after potency prediction, novelty filtering and practical selection. They were not a random sample of the original 40.6 million, and the experiment therefore does not show that APEX labels every active and inactive peptide with 91 percent accuracy.
Nor are the 53 laboratory hits “new antibiotics” in the clinical sense. They are early discovery-stage leads. A compound that inhibits bacteria in broth must still survive tests of toxicity, stability, distribution, metabolism, dosing, manufacturing, resistance and efficacy in living organisms.
The peptides appear to attack a bacterium’s electrical boundary
The model did not only find unfamiliar sequences. Many candidates shared physical features that could explain their activity. They carried a high net positive charge and were relatively hydrophobic.
Bacterial membranes are rich in negatively charged components. Positive charge can draw a peptide toward that surface, while hydrophobic regions can interact with the membrane’s lipid interior. Circular-dichroism experiments showed that many VEPs shifted from flexible structures into alpha helices in membrane-mimicking environments, a shape change consistent with membrane activity.
The researchers then watched what selected peptides did to bacterial membranes with fluorescent probes. Twenty-three permeabilized the outer membrane of P. aeruginosa. Among 28 tested for effects on that bacterium’s cytoplasmic membrane, 26 caused stronger depolarization than the polymyxin B and levofloxacin control groups in that assay.
Depolarization collapses the voltage difference a bacterium uses for essential cellular work. The team concluded that cytoplasmic-membrane depolarization was the candidates’ main observed antimicrobial mechanism. That resembles the action of other antimicrobial peptides, even though the new sequences occupy less familiar territory.
ScienceBlog previously covered an AI screen that identified halicin in a chemical library. That molecule and these venom fragments are chemically different, but the logic is shared: computation ranks a search space, while laboratory experiments determine whether the ranking corresponds to biological activity.
Three candidates reached a small mouse experiment
The researchers selected three active peptides with comparatively favorable toxicity profiles. They came from proteins associated with a scorpion, a cone snail and a wolf spider. Each was tested in a skin-wound model using A. baumannii, a Gram-negative pathogen notorious for hospital-acquired and drug-resistant infections.
Six mice were used per group. Two hours after infection, the investigators applied one topical dose of a peptide at its MIC. All three reduced bacterial counts after two days. The strongest, Arachnoserver-5, produced about a two-log reduction, roughly 100-fold, relative to untreated controls. By day four it produced about a three-log, or 1,000-fold, reduction and performed at a level comparable to the antibiotic controls in that model.
No significant body-weight change was observed in the treated mice. That is encouraging, but “without observable toxicity” in this experiment has narrow meaning. It refers to a single topical treatment, this model, these group sizes and the reported measures. It does not establish safe intravenous or oral dosing, prolonged use, reproductive safety or safety in humans.
The danger in venom did not disappear
Potency and toxicity can arise from the same physical properties. A strongly charged, hydrophobic peptide may disrupt a bacterium, but it can also harm mammalian cells.
Some VEPs damaged cultured human embryonic kidney cells. A few spider-database candidates also ruptured human red blood cells and were excluded from the mouse work. The authors reported that nearly 40 percent of the identified peptides were not predicted to affect potassium channels. Turned around, about 60 percent, including the peptides advanced into mice, were predicted to modulate ion channels.
That is not a minor detail. Venoms evolved partly because their molecules interfere efficiently with animal physiology, and ion channels are frequent targets. The authors called for direct electrophysiological tests and further chemical optimization, especially if anyone proposes systemic rather than topical use.
Peptides also face ordinary drug-development obstacles. Enzymes can break them down. They may clear too quickly, fail to reach an infection at useful concentrations or provoke unwanted immune effects. Stability, bioavailability and selectivity all remain optimization problems.
AI accelerated triage, not the rest of medicine
The Penn result fits a growing style of antibiotic discovery. Researchers treat biological sequence databases as searchable chemical archives, then use models to decide which tiny fraction is worth manufacturing. A previous ScienceBlog report described how another team mined the human gut microbiome for antimicrobial peptides. Venom supplies a different archive, shaped by millions of years of molecular competition and predation.
The honest promise is not that AI found 53 pharmacy-ready antibiotics in an afternoon. It found 53 laboratory-active molecules among 58 carefully chosen tests after screening a virtual space no laboratory could exhaust molecule by molecule.
That changes where human effort begins. Chemists and microbiologists can spend time on dozens of informed candidates rather than millions of unranked possibilities. The long work after a hit remains long, but choosing better experiments is itself a consequential scientific advance.