Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • Sarrus
    Junior Member
    • Oct 2023
    • 1

    #1

    ONT sequencing methylation data normalisation

    I have methylation data from ONT sequencing expressed as percentage of methylation and/or number of reads. I don't know if I'm supposed to normalize the data and how to handle replicates. Some ideas: I have data in triplicate. I want to normalize the data (quantile normalization is the preferred option) and/or collapse data to have 1 file/dataset from 3 replicates (doing weighted average maybe). So 1 idea would be just to perform weighted average of each methylation value doing [%M*NR] where %M is the percentage of methylated reads and NR is number of reads, per each methylated position. In this case I will lose the meaning of the percentage value in the downstream analyses because I will have a weighted value. otherwise I can perform quantile normalization on the number of reads (total and methylated) and then calculate percentage of methylated reads. I would like the opinion of someone with a better statistical background than me thanks for your kind help!

    %M-S1 NR-S1 %M-S2 NR-S2 %M-S3 NR-S3
    Meth1 20 60 15 54 41 12
    Meth2 40 14 78 52 13 65
    Meth3 12 94 73 19 37 70
    Meth4 36 77 69 14 26 74

    0
    Bioinformatics
    0%
    0
    Oxford Nanopore
    0%
    0
    Methylation
    0%
    0
  • BasesBreaker
    Junior Member
    • Aug 2023
    • 5

    #2
    Okay, so here's my take on this. Since you have triplicates and want to both normalize the data and condense it, this is my recommendation.

    Start by normalizing the number of reads (both total and methylated) across your replicates using quantile normalization. In this case, you'll ensure that all your samples have a similar distribution, which should make it easier to compare them. Now after normalization, you should recalculate the percentage of methylated reads for each position.

    And if you want to collapse the data from 3 replicates into 1 dataset, then a weighted average makes the most sense to me. For each position, you can calculate the weighted methylation percentage using the formula you've included in your post. After you get this weighted value for each replicate, average these values for the three replicates. This will end up giving you a single value that takes into account both the methylation percentage and the number of reads.

    If you do all of this, you will normalize the data and then condense it into a single dataset that's more representative of your three replicates.

    I'm a bit rusty at this type of work so it wouldn't hurt to get a second opinion, but based on what I remember, that's what I would do.​

    Comment

    Latest Articles

    Collapse

    • SEQadmin2
      Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
      by SEQadmin2



      CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

      Despite this, “CRISPR helped turn genome editing from a specialized technique into
      ...
      07-31-2026, 11:01 AM
    • SEQadmin2
      Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
      by SEQadmin2


      Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

      The systematic characterization of the human proteome has
      ...
      07-20-2026, 11:48 AM
    • SEQadmin2
      Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
      by SEQadmin2



      Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
      ...
      07-09-2026, 11:10 AM

    ad_right_rmr

    Collapse

    News

    Collapse

    Topics Statistics Last Post
    Started by SEQadmin2, Yesterday, 10:13 AM
    0 responses
    14 views
    0 reactions
    Last Post SEQadmin2  
    Started by SEQadmin2, 07-31-2026, 02:55 AM
    0 responses
    29 views
    0 reactions
    Last Post SEQadmin2  
    Started by SEQadmin2, 07-24-2026, 12:17 PM
    0 responses
    23 views
    0 reactions
    Last Post SEQadmin2  
    Started by SEQadmin2, 07-23-2026, 11:41 AM
    0 responses
    21 views
    0 reactions
    Last Post SEQadmin2  
    Working...