Unconfigured Ad

Collapse

Pathogen Surveillance with Advanced Genomic Tools

Collapse
X
Collapse
  •  

  • Pathogen Surveillance with Advanced Genomic Tools

    Click image for larger version  Name:	Pathogen Article Image2.jpg Views:	0 Size:	430.9 KB ID:	326508




    The COVID-19 pandemic highlighted the need for proactive pathogen surveillance systems. As ongoing threats like avian influenza and newly emerging infections continue to pose risks, researchers are working to improve how quickly and accurately pathogens can be identified and tracked. In a recent SEQanswers webinar, two experts discussed how next-generation sequencing (NGS) and machine learning are shaping efforts to monitor viral variation and trace the origins of infectious outbreaks.

    Modern Pathogen Surveillance through NGS
    In the opening presentation, Sion Bayliss, Ph.D., Lecturer in Endemics, Epizootics, and Emerging Infectious Diseases at Bristol Veterinary School, discussed his work on developing a machine learning framework that classifies the geographic origin of Salmonella enterica serovar Enteritidis using genomic surveillance data1. S. Enteritidis causes various types of infection, some of which can be severe and systemic, and is primarily transmitted through the consumption of contaminated food, particularly poultry, meat, and eggs.

    Identifying where infections originate geographically is a key aspect of pathogen surveillance. With over 8,500 cases of Salmonella and 2,500 cases of this subspecies sequenced annually by the UK Health Security Agency (UKHSA), this method offers an efficient way to support outbreak investigations.

    Bayliss and his colleagues aimed to improve outbreak responses by accurately predicting the geographical origin of S. Enteritidis and rapidly providing detailed information for epidemiologists. To do this, they used genomic data from isolates and reported travel history collected between 2014 and 2019. While around 10,000 genomes were available, only 3,000 included the necessary metadata to train the model.

    The group first observed that S. Enteritidis exhibits a strong phylogeographic signal, with genetically related isolates often originating from the same regions. This supported the use of machine learning to predict geographic origin based on genomic data. They implemented a hierarchical classification approach, training a random forest model to assign isolates by continent, subregion, and country. Each level included confidence scores, allowing for more detailed information when possible while flagging uncertain predictions. Bayliss noted that the models were optimized for speed that processed raw sequence data in under four minutes and used by used unitigs derived from de novo graph assemblies that captured both core SNPs and accessory gene content.

    However, several challenges related to sampling bias emerged. Some countries were poorly represented and genetically diverse which limited model performance. The skewed sampling reflected where UK residents travel as well as varying infection risks and sequencing efforts abroad. To overcome these imbalances, the team applied resampling techniques and tested a range of model-resampler combinations before deciding on the final framework.

    While sharing some of the results, Bayliss noted that the hierarchical structure allows for clear visualization of confidence levels across regions, with strong performance in areas that had ample, genetically consistent data. The model accuracy was highest at the continental level but declined with more granular predictions. Despite this, performance remained stable over time, suggesting that retraining every two to three years is sufficient. Additionally, external validation with outbreak datasets showed strong predictive performance, even with varying strain diversity.

    Even with strong results, Bayliss acknowledged a limitation that emerged during the analysis of a Polish egg outbreak. Although the source was known, the model consistently misclassified the isolates because many affected individuals had traveled to other countries where they consumed the contaminated product, skewing the metadata. This underscored a key issue that country-level travel data doesn’t always reflect the true origin of infection, especially in a globally interconnected food system.

    Bayliss concluded his presentation by explaining that hierarchical machine learning models can rapidly and accurately predict the geographic origin of Salmonella Enteritidis and other serovars, even when new or temporally variable data are introduced. Since the models depend on the regions present in the training data, they can misinterpret samples when metadata is inaccurate or incomplete. Ongoing work in his research group aims to address these gaps by integrating temporal data, external datasets, and deep learning approaches.

    Genomic Surveillance of Rhinoviruses
    In the second presentation, Stephanie Goya, Ph.D., Postdoctoral Research Scientist in the Greninger Lab at the University of Washington, shared her work on the genomic surveillance of rhinoviruses2. She began by describing the burden of rhinoviruses, which are the primary cause of the common cold and exacerbate conditions like asthma, cystic fibrosis, and chronic obstructive pulmonary disease. Despite being one of the most common respiratory viruses worldwide, rhinoviruses remain heavily understudied.

    During the COVID-19 pandemic, global containment efforts led to a sharp decline in the circulation of most respiratory viruses. However, rhinoviruses continued to widely circulate, which prompted Goya and her team to investigate their genetic characteristics and the potential emergence of new variants. Previous genomic surveillance of rhinoviruses had relied almost entirely on partial sequencing of the VP1 region, with fewer than 500 complete genomes publicly available across all three species (A, B, and C). To address this knowledge gap, the team designed a study using metagenomics to capture full genomes from samples collected at COVID-19 testing sites in western Washington between 2021 and 2023.

    Rhinovirus-positive samples were screened by qPCR, and only those with CT values below 33 were selected to ensure sufficient viral load for successful metagenomic sequencing. Using Illumina technology and in-house software called REVICA, the team generated over 1,000 complete rhinovirus genomes. This more than doubled the global collection of full-length sequences.

    A key finding was that all three rhinovirus species circulated across the region, with rhinovirus C more common in winter and rhinovirus A in summer. Surprisingly, the team detected over half of all known rhinovirus genotypes within this single geographic area, and many genotypes co-circulated over time. Some of these lineages trace back decades, with estimated common ancestors as early as the 1980s.
    Even with their long circulation, many genotypes remain well-conserved at the amino acid level. The study also identified evidence of recombination between genotypes, including a breakpoint in the 3C protease. Goya emphasized that these findings were only possible through full-genome sequencing, as conventional methods focused only on VP1 would not have detected the recombination event.

    In closing, Goya noted that although partial sequencing supports routine epidemiological tracking, full-genome surveillance provides a more complete understanding of viral variation and persistence. Although scaling up these efforts requires addressing key methodological challenges. Goya also shared that metagenomics is powerful but limited to high-viral-load samples, while amplicon sequencing is more cost-effective but depends on prior genomic knowledge, and hybridization capture is thorough yet expensive. Lastly, Goya emphasized that standardized molecular epidemiology frameworks and broader community collaboration are needed to fill existing data gaps and strengthen genomic surveillance of rhinoviruses and other understudied pathogens.

    References
    1. Sion C Bayliss, Rebecca K Locke, Claire Jenkins, Marie Anne Chattaway, Timothy J Dallman, Lauren A Cowley (2023) Rapid geographical source attribution of Salmonella enterica serovar Enteritidis genomes using hierarchical machine learning eLife 12:e84167 https://doi.org/10.7554/eLife.84167
    2. Stephanie Goya, Seffir T Wendm, Hong Xie, Tien V Nguyen, Sarina Barnes, Rohit R Shankar, Jaydee Sereewit, Kurtis Cruz, Ailyn C Pérez-Osorio, Margaret G Mills, Alexander L Greninger, Genomic Epidemiology and Evolution of Rhinovirus in Western Washington State, 2021–2022, The Journal of Infectious Diseases, Volume 231, Issue 1, 15 January 2025, Pages e154–e164, https://doi.org/10.1093/infdis/jiae347
      Please sign into your account to post comments.

    About the Author

    Collapse

    seqadmin Benjamin Atha holds a B.A. in biology from Hood College and an M.S. in biological sciences from Towson University. With over 9 years of hands-on laboratory experience, he's well-versed in next-generation sequencing systems. Ben is currently the editor for SEQanswers. Find out more about seqadmin

    Latest Articles

    Collapse

    • New Genomics Technologies Take Aim at Long-Standing Limits
      by SEQadmin2


      Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

      We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
      ...
      09-28-2026, 10:25 AM
    • How Immunogenomics Decodes Immunity’s Genetic Blueprint
      by SEQadmin2




      The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

      This convergence of genetics, immunology, and computation...
      09-01-2026, 05:41 AM
    • Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
      by SEQadmin2



      CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

      Despite this, “CRISPR helped turn genome editing from a specialized technique into
      ...
      07-31-2026, 11:01 AM

    ad_right_rmr

    Collapse

    News

    Collapse

    Working...