Data availability
The data supporting the results of this study are available within the article and its Supplementary Information. The training data are provided in Supplementary Information, referred to as ‘general classifier’. Source data are provided with this paper.
Code availability
The custom machine learning codes and training data for the identification of amino acids, NMPs, saccharides and peptides, as well as for few-shot learning, are provided in Supplementary Information.
References
-
Aebersold, R. & Mann, M. Mass-spectrometric exploration of proteome structure and function. Nature 537, 347–355 (2016).
-
Aebersold, R. & Mann, M. Mass spectrometry-based proteomics. Nature 422, 198–207 (2003).
-
Cravatt, B. F., Simon, G. M. & Yates Iii, J. R. The biological impact of mass-spectrometry-based proteomics. Nature 450, 991–1000 (2007).
-
Nesvizhskii, A. I. A survey of computational methods and error rate estimation procedures for peptide and protein identification in shotgun proteomics. J. Proteomics 73, 2092–2123 (2010).
-
Deol, H. et al. After 75 years, an alternative to Edman degradation: a mechanistic and efficiency study of a base-induced method for N-terminal peptide sequencing. J. Am. Chem. Soc. 147, 13973–13982 (2025).
-
Holley, R. W. et al. Structure of a ribonucleic acid. Science 147, 1462–1465 (1965).
-
Sedlazeck, F. J., Lee, H., Darby, C. A. & Schatz, M. C. Piercing the dark matter: bioinformatics of long-range sequencing and mapping. Nat. Rev. Genet. 19, 329–346 (2018).
-
Abdel-Ghany, S. E. et al. A survey of the sorghum transcriptome using single-molecule long reads. Nat. Commun. 7, 11706 (2016).
-
Zhao, H. et al. Chromosome-level reference genome and alternative splicing atlas of moso bamboo (Phyllostachys edulis). GigaScience 7, giy115 (2018).
-
Li, S., Yamada, M., Han, X., Ohler, U. & Benfey, P. N. High-resolution expression map of the Arabidopsis root reveals alternative splicing and lincRNA regulation. Dev. Cell 39, 508–522 (2016).
-
Weirather, J. L. et al. Characterization of fusion genes and the significantly expressed fusion isoforms in breast cancer by hybrid sequencing. Nucleic Acids Res. 43, e116 (2015).
-
Garalde, D. R. et al. Highly parallel direct RNA sequencing on an array of nanopores. Nat. Methods 15, 201–206 (2018).
-
Zhao, L. et al. Analysis of transcriptome and epitranscriptome in plants using PacBio Iso-Seq and nanopore-based direct RNA sequencing. Front. Genet. 10, 253 (2019).
-
Urban, J. et al. Predicting glycan structure from tandem mass spectrometry via deep learning. Nat. Methods 21, 1206–1215 (2024).
-
Yao, H.-Y. -Y., Wang, J.-Q., Yin, J.-Y., Nie, S.-P. & Xie, M.-Y. A review of NMR analysis in polysaccharide structure and conformation: progress, challenge and perspective. Food Res. Int. 143, 110290 (2021).
-
Krishnamoorthy, L. & Mahal, L. K. Glycomic analysis: an array of technologies. ACS Chem. Biol. 4, 715–732 (2009).
-
Manrao, E. A. et al. Reading DNA at single-nucleotide resolution with a mutant MspA nanopore and phi29 DNA polymerase. Nat. Biotechnol. 30, 349–353 (2012).
-
Nivala, J., Marks, D. B. & Akeson, M. Unfoldase-mediated protein translocation through an α-hemolysin nanopore. Nat. Biotechnol. 31, 247–250 (2013).
-
Yu, L. et al. Unidirectional single-file transport of full-length proteins through a nanopore. Nat. Biotechnol. 41, 1130–1139 (2023).
-
Sauciuc, A., Morozzo della Rocca, B., Tadema, M. J., Chinappi, M. & Maglia, G. Translocation of linearized full-length proteins through an engineered nanopore under opposing electrophoretic force. Nat. Biotechnol. 42, 1275–1281 (2024).
-
Li, M. et al. Identification of tagged glycans with a protein nanopore. Nat. Commun. 14, 1737 (2023).
-
Xia, B. et al. Mapping the acetylamino and carboxyl groups on glycans by engineered α-hemolysin nanopores. J. Am. Chem. Soc. 145, 18812–18824 (2023).
-
Brinkerhoff, H., Kang, A. S. W., Liu, J., Aksimentiev, A. & Dekker, C. Multiple rereads of single proteins at single-amino acid resolution using nanopores. Science 374, 1509–1513 (2021).
-
Yan, S. et al. Single molecule ratcheting motion of peptides in a Mycobacterium smegmatis porin A (MspA) nanopore. Nano Lett. 21, 6703–6710 (2021).
-
Motone, K. et al. Multi-pass, single-molecule nanopore reading of long protein strands. Nature 633, 662–669 (2024).
-
Bonini, A., Sauciuc, A. & Maglia, G. Engineered nanopores for exopeptidase protein sequencing. Nat. Methods 21, 16–17 (2024).
-
Haworth, W. N., Peat, S. & Bourne, E. J. Synthesis of amylopectin. Nature 154, 236–236 (1944).
-
Meyer, K. H., Gürtler, P. & Bernfeld, P. Structure of amylopectin. Nature 160, 900–901 (1947).
-
Faller, M., Niederweis, M. & Schulz, G. E. The structure of a mycobacterial outer-membrane channel. Science 303, 1189–1192 (2004).
-
Cao, J. et al. Giant single molecule chemistry events observed from a tetrachloroaurate(III) embedded Mycobacterium smegmatis porin A nanopore. Nat. Commun. 10, 5668 (2019).
-
Zhang, S. et al. A nanopore-based saccharide sensor. Angew. Chem. Int. Ed. Engl. 61, e202203769 (2022).
-
Zhang, S. et al. Discrimination of disaccharide isomers of different glycosidic linkages using a modified MspA nanopore. Angew. Chem. Int. Ed. Engl. 63, e202316766 (2024).
-
Wang, Y. et al. Identification of nucleoside monophosphates and their epigenetic modifications using an engineered nanopore. Nat. Nanotechnol. 17, 976–983 (2022).
-
Wang, K. et al. Unambiguous discrimination of all 20 proteinogenic amino acids and their modifications by nanopore. Nat. Methods 21, 92–101 (2024).
-
Sun, W. et al. Nanopore discrimination of rare earth elements. Nat. Nanotechnol. 20, 523–531 (2025).
-
Cal, P. M. S. D. et al. Iminoboronates: a new strategy for reversible protein modification. J. Am. Chem. Soc. 134, 10299–10305 (2012).
-
Cambray, S. & Gao, J. Versatile bioconjugation chemistries of ortho-boronyl aryl ketones and aldehydes. Acc. Chem. Res. 51, 2198–2206 (2018).
-
Haggett, J. G. & Domaille, D. W. ortho-Boronic acid carbonyl compounds and their applications in chemical biology. Chemistry 30, e202302485 (2024).
-
Arnal-Hérault, C. et al. Functional G-quartet macroscopic membrane films. Angew. Chem. Int. Ed. Engl. 46, 8409–8413 (2007).
-
Hutin, M., Bernardinelli, G. & Nitschke, J. R. An iminoboronate construction set for subcomponent self-assembly. Chemistry 14, 4585–4593 (2008).
-
Akgun, B. & Hall, D. G. Boronic acids as bioorthogonal probes for site-selective labeling of proteins. Angew. Chem. Int. Ed. Engl. 57, 13028–13044 (2018).
-
Breiman, L. Random forests. Mach. Learn. 45, 5–32 (2001).
-
Fisher, A., Rudin, C. & Dominici, F. All models are wrong, but many are useful: learning a variable’s importance by studying an entire class of prediction models simultaneously. J. Mach. Learn. Res. 20, 177 (2019).
-
Boersma, A. J. & Bayley, H. Continuous stochastic detection of amino acid enantiomers with a protein nanopore. Angew. Chem. Int. Ed. Engl. 51, 9606–9609 (2012).
-
Guo, Y. et al. Metal–organic complex-functionalized protein nanopore sensor for aromatic amino acids chiral recognition. Analyst 142, 1048–1053 (2017).
-
Zhang, M. et al. Real-time detection of 20 amino acids and discrimination of pathologically relevant peptides with functionalized nanopore. Nat. Methods 21, 609–618 (2024).
-
Ayub, M., Hardwick, S. W., Luisi, B. F. & Bayley, H. Nanopore-based identification of individual nucleotides for direct RNA sequencing. Nano Lett. 13, 6144–6150 (2013).
-
Ramsay, W. J. & Bayley, H. Single-molecule determination of the isomers of d-glucose and d-fructose that bind to boronic acids. Angew. Chem. Int. Ed. Engl. 57, 2841–2845 (2018).
-
Ouldali, H. et al. Electrical recognition of the twenty proteinogenic amino acids using an aerolysin nanopore. Nat. Biotechnol. 38, 176–181 (2020).
-
Zhang, Y. et al. Peptide sequencing based on host–guest interaction-assisted nanopore sensing. Nat. Methods 21, 102–109 (2024).
-
Piguet, F. et al. Identification of single amino acid differences in uniformly charged homopolymeric peptides with aerolysin nanopore. Nat. Commun. 9, 966 (2018).
-
Singh, P. R. et al. Pulling peptides across nanochannels: resolving peptide binding and translocation through the hetero-oligomeric channel from Nocardia farcinica. ACS Nano 6, 10699–10707 (2012).
-
Ji, Z., Kang, X., Wang, S. & Guo, P. Nano-channel of viral DNA packaging motor as single pore to differentiate peptides with single amino acid difference. Biomaterials 182, 227–233 (2018).
-
Miyagi, M., Takiguchi, S., Hakamada, K., Yohda, M. & Kawano, R. Single polypeptide detection using a translocon EXP2 nanopore. Proteomics 22, e2100070 (2022).
-
Huang, G., Voet, A. & Maglia, G. FraC nanopores with adjustable diameter identify the mass of opposite-charge peptides with 44 dalton resolution. Nat. Commun. 10, 835 (2019).
-
Ratinho, L., Bacri, L., Thiebot, B., Cressiot, B. & Pelta, J. Identification and detection of a peptide biomarker and its enantiomer by nanopore. ACS Cent. Sci. 10, 1167–1178 (2024).
-
Wei, X., Wen, J., Wu, H., Qu, Z. & Huang, G. Obtaining narrow distributions of single-molecule peptide signals enables sensitive peptide discrimination with α-hemolysin nanopores. J. Am. Chem. Soc. 147, 9304–9315 (2025).
-
Wang, Y., Yao, Q., Kwok, J. T. & Ni, L. M. Generalizing from a few examples: a survey on few-shot learning. ACM Comput. Surv. 53, 63 (2020).
-
Tomé, D. Yeast extracts: nutritional and flavoring food ingredients. ACS Food Sci. Technol. 1, 487–494 (2021).
-
Tao, Z. et al. Yeast extract: characteristics, production, applications and future perspectives. J. Microbiol. Biotechnol. 33, 151–166 (2023).
-
He, M., Zhou, X. & Wang, X. Glycosylation: mechanisms, biological functions and clinical implications. Signal Transduct. Target. Ther. 9, 194 (2024).
Funding
This project was funded by the National Key R&D Program of China (2023YFF1205900 and 2022YFA1304602), National Natural Science Foundation of China (22225405 and 22534004), State Key Laboratory of Analytical Chemistry for Life Science (5431ZZXM2509), Yachen Foundation of Nanjing University, National Natural Science Foundation of China (223B2402, to K.F.W). China National Postdoctoral Program for Innovative Talents (BX20250087, to K.F.W). Jiangsu Funding Program for Excellent Postdoctoral Talent (2025ZB212, to K.F.W) and China Postdoctoral Science Foundation (2025M780957, to K.F.W).
Ethics declarations
Competing interests
S.H. and L.Y. have filed patents describing the preparation of MspA containing a single FPBA adaptor and its applications thereof. The other authors declare no other competing interests.
Peer review
Peer review information
Nature Biotechnology thanks Manish Kumar who co-reviewed with Priyanshu R. Gupta, Meni Wanunu and the other, anonymous reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data
Extended Data Fig. 1 Amino acid, saccharide and NMP sensing performed by different types of nanopores.
(a) Sensing of phenylalanine (Phe) with MspA-PBA. Left: The structure of MspA-PBA. Right: A representative trace of Phe sensing performed with MspA-PBA. (b) Sensing of D-fructose (Fru) with MspA-NTA-Ni. Left: The structure of MspA-NTA-Ni. Right: A representative trace of Fru sensing performed with MspA-NTA-Ni. (c) Sensing of guanosine 5’-monophosphate (GMP) with MspA-NTA-Ni. Left: The structure of MspA-NTA-Ni. Right: A representative trace of GMP sensing performed with MspA-NTA-Ni. The final concentrations of Phe, Fru and GMP were 400 μM, 20 mM and 2 mM, respectively. No nanopore events were observed in all above measurements. The open pore current of MspA-PBA ( + 160 mV bias) and MspA-NTA-Ni (+100 mV bias) were defined as IPBA and INTA-Ni, respectively. The corresponding open pore current was marked by a gray dashed line.
Extended Data Fig. 2 Distinguishing 21 proteinogenic amino acids, 4 canonical NMPs and 4 saccharides using MspA-FPBA.
(a) Representative events of different amino acids, NMPs and saccharides acquired by MspA-FPBA (Methods). 21 amino acids (including 20 common proteinogenic amino acids and selenocysteine), 4 NMPs (adenosine 5’-monophosphate, AMP; uridine 5’-monophosphate, UMP; cytidine 5’-monophosphate, CMP; guanosine 5’-monophosphate, GMP), and 4 saccharides (iduronic acid, IdoA; L-arabinose, Ara; D-fructose, Fru; N-acetyl-D-glucosamine, GlcNAc) were respectively added to the cis side during the measurement. Cysteine, histidine and lysine each have two representative types of events, respectively labeled as C1/2, H1/2 and K1/2. Notably, cysteine was blocked with maleimide prior to the measurement to avoid the appearance of long-residing events (Supplementary Fig. 14). Each NMPs has two representative types of events, respectively labeled as AMP1/2, UMP1/2, CMP1/2, and GMP1/2. Ara has three types of representative events, labeled as Ara1/2/3. Fru and GlcNAc both have two types of representative events, labeled as Fru1/2 and GlcNAc1/2, respectively. (b) The scatter plot of ΔI versus SD for events of the 21 amino acids. The scatter plot was generated using at least 100 events for each amino acid. (c) The scatter plot of ΔI versus SD for the type 1 events (A/U/C/GMP1) of 4 NMPs. At least 100 type 1 events of each NMP were used to generate the plot. (d) The scatter plot of ΔI versus SD for the representative events of the 4 saccharides. The plot was generated using at least 100 events for each saccharide.
Extended Data Fig. 3 The scatter plot of ΔI versus SD of events acquired with different amino acids.
At least 100 events acquired with each amino acid were used to generate the plot, according to which, most amino acid events are fully distinguishable. To clarify the detail, the events inside the red box are further zoomed-in and shown on the right. Although the events corresponding to P and M appear to overlap in the plot, their event characteristics are visually different and can be discriminated in the 3D event scatter plot of ΔI, SD, and max (Supplementary Fig. 18). High-precision identification of P and M events can be achieved by machine learning, with validation accuracies of 98.8% and 99.4% respectively (Fig. 3a).
Extended Data Fig. 4 Type 1 and type 2 events of NMPs.
The measurements were carried out as described in Methods. A 1.5 M KCl buffer (cis: 1.5 M KCl, 10 mM CHES, pH 9.0; trans: 1.5 M KCl, 10 mM MES, pH 6.0) was used. A transmembrane voltage of +160 mV was continually applied. (a-d) Left: Chemical structures of nucleoside monophosphates. Right: Representative traces containing events caused by AMP (a), UMP (b), CMP (c) and GMP (d) binding. Each NMP was separately added to cis with a final concentration of 2 mM. IFPBA represents the open pore current of MspA-FPBA. Each nucleoside monophosphate generates type 1 events (AMP1, UMP1, CMP1 and GMP1) and type 2 events (AMP2, UMP2, CMP2 and GMP2). A/U/C/GMP1 + 2 represents the combination of type 1 events and type 2 events. (e) The scatter plot of ΔI versus SD generated by type 1 events (marked with a red dashed box) and type 2 events (marked with a gray dashed ellipse) of nucleoside monophosphates. Type 1 events (AMP1, n = 113; UMP1, n = 207; CMP1, n = 169; GMP1, n = 190) and type 2 events (AMP2, n = 125; UMP2, n = 437; CMP2, n = 388; GMP2, n = 178) were employed to generate the statistics. (f) Top: Representative type 1 events of the corresponding nucleoside monophosphates. Bottom: Representative type 2 events of the corresponding nucleoside monophosphates.
Extended Data Fig. 5 Identification of amino acids with PTMs and epigenetic NMPs by MspA-FPBA.
(a) The chemical structures of amino acids and NMPs with modifications. N6-Acetyl-L-lysine (Ac-K), O-Phospho-L-tyrosine (P-Y), and N3,N4-Dimethyl-L-arginine (Me-R) represent the acetylation, phosphorylation and methylation modifications of amino acids, respectively. Three epigenetic NMPs, namely pseudouridine-5’-monophosphate (ψ), N6-Methyladenosine-5’-monophosphate (m6A), and inosine-5’-monophosphate (IMP) were also investigated. (b) Representative traces acquired during MspA-FPBA sensing of Ac-K, P-Y, and Me-R, respectively. Each amino acid was added to the cis side, reaching a final concentration of 400 μM. The IFPBA was marked with a gray dashed line. (c) The scatter plot of ΔI versus SD for the Ac-K, P-Y, and Me-R events obtained as described in b. At least 100 events of each analyte were used to generate the scatter plot. (d) Representative traces obtained during MspA-FPBA sensing of ψ, m6A and IMP. Each analyte was added to the cis side at a final concentration of 2 mM. IFPBA was marked with a gray dashed line. (e) The scatter plot of ΔI versus SD for the ψ1, m6A1 and IMP1 events obtained from d. At least 100 events of each analyte were used to generate the scatter plot.
Extended Data Fig. 6 Machine learning identification of 24 amino acids, 7 nucleoside monophosphates, 4 saccharides, and 5 peptides.
(a) The workflow of machine learning. Nanopore events acquired with 24 amino acids (21 proteinogenic amino acids and 3 amino acids with post translational modifications), 7 nucleoside monophosphates (4 canonical nucleoside monophosphates and 3 nucleoside monophosphates with epigenetic modifications), 4 saccharides, and 5 peptides (IK, LRG, SLR, GYIK, and LRGY) were collected to form the training dataset. Nine event features, including ΔI, SD, LocalSD Ratio, min, max, range, toff, skew and kurt, were employed to form a nine-dimensional feature matrix. All models were evaluated with the 10-fold cross-validation accuracies, according to which Ensemble model (bagged trees) has performed the best with a validation accuracy of 98.7% (Supplementary Table 7). (b) The confusion matrix for amino acids (orange), nucleoside monophosphates (blue), saccharides (green), nucleoside monophosphates with epigenetic modifications (brown), amino acids with post translational modifications (pink), and peptides (purple) generated by the bagged trees model employing a nine-dimensional feature matrix. The row of the matrix represents the true class and the column represents the predicted class.
Extended Data Fig. 7 Identification of amino acids and glycosylation modifications by protease digestion of glycosylated peptides.
(a) Schematic diagram of the identification of amino acids and glycosylation modifications following the enzymatic cleavage of glycosylated peptides using MspA-FPBA. Leucine aminopeptidase (LAP) was employed to cleave the N-terminal amino acids of glycosylated peptides until reaching the glycosylation site that impeded enzymatic digestion, thereby exposing the glycosylation site. Then, MspA-FPBA was used to identify the generated amino acids, thereby confirming the amino acid components before the glycosylation site. For the short peptide containing the glycosylation site, both the N-terminal amino group and the glycosyl group could be recognized by MspA-FPBA, resulting in multiple types of events. (b) Representative traces of the glycosylated heptapeptide with the sequence GMQRN(GlcNAc)IS following LAP enzymatic digestion (Methods). Representative events corresponding to different amino acids were automatically predicted via machine learning and labeled with color-coded dots (G: dark red; M: dark purple; Q: light green; R: orange). Three unidentified nanopore events (GP1/2/3) were attributed to the N(GlcNAc)IS peptide generated by enzymatic digestion, with the glycosylation site being exposed. (c) The scatter plot of ΔI versus SD for the amino acid and glycosylated tripeptide events obtained from (b). The scatter plot was generated using a 10 min continually recorded trace. The four clusters of amino acids corresponding to G, M, Q, and R were identified by machine learning. The three distinct clusters corresponding to the N(GlcNAc)IS peptide were respectively labeled as GP1, GP2 and GP3 based on the blockage amplitude and noise amplitude. All measurements (b–c) were conducted under a continuous bias of +160 mV, with 25 µL of the enzymatic hydrolysate added to the cis side.
Supplementary information
Supplementary Information (download PDF )
Supplementary Tables 1–18, Figs. 1–54 and Materials.
Reporting Summary (download PDF )
Peer Review File (download PDF )
Supplementary Video 1 (download MP4 )
Simultaneous sensing of Phe, GMP and Fru. The measurement was performed using MspA-FPBA in a 1.5 M KCl buffer (cis: 1.5 M KCl and 10 mM CHES pH 9.0; trans: 1.5 M KCl and 10 mM MES pH 6.0). A potential of +160 mV was continually applied. Phe, GMP and Fru were simultaneously added to the cis side at final concentrations of 50 μM, 1 mM and 5 mM, respectively. On the basis of the blockage amplitude and SD of the corresponding standard substances, each event was identified and labeled with color-coded dots (Phe, orange; GMP1/2, blue; Fru1/2, green).
Supplementary Video 2 (download MP4 )
Identification of amino acids and glycosylation modifications from the glycosylated heptapeptide Gly-Met-Gln-Arg-Asn-(GlcNAc)-Ile-Ser following LAP enzymatic digestion. The measurement was performed using MspA-FPBA in a 1.5 M KCl buffer (cis: 1.5 M KCl and 10 mM CHES pH 9.0; trans: 1.5 M KCl and 10 mM MES pH 6.0). A potential of +160 mV was continually applied, with 25 µl of the enzymatic hydrolysate added to the cis side. With a trained machine learning algorithm, each amino acid event was automatically identified and labeled with color-coded dots (Gly, dark red; Met, dark purple; Gln, light green; Arg, orange). Three unidentified nanopore events were respectively labeled as GP1, GP2 and GP3 based on the blockage amplitude. GP1, GP2 and GP3 events were attributed to the N(GlcNAc)IS peptide generated by enzymatic digestion, with the glycosylation site being exposed.
Supplementary Video 3 (download MP4 )
Identification of amino acids, NMPs, saccharides and peptides from yeast cell extract. The measurement was performed using MspA-FPBA in a 1.5 M KCl buffer (cis: 1.5 M KCl and 10 mM CHES pH 9.0; trans: 1.5 M KCl and 10 mM MES pH 6.0). A potential of +160 mV was continually applied, with 5 µl of the yeast cell extract filtrate added to the cis side. Various nanopore events were automatically predicted by few-shot learning and labeled with color-coded dots (amino acids, yellow; NMPs, blue; saccharides, green; peptides, purple). With the trained bagged trees model, different amino acids and NMPs were further identified.
Supplementary Software 1
The machine learning source code and training data.
Source data
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Yao, L., Wang, Z., Chen, J. et al. An engineered nanopore identifies saccharides, amino acids, peptides and ribonucleotides.
Nat Biotechnol (2026). https://doi.org/10.1038/s41587-026-03308-9
-
Received:
-
Accepted:
-
Published:
-
Version of record:
-
DOI: https://doi.org/10.1038/s41587-026-03308-9