Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • wynstep
    Member
    • Jan 2014
    • 11

    #1

    miRNA Illumina sequencing - low alignment rate

    Dear Colleagues,
    I want to share my strange experience with you, to ask your opinions and help.

    I'm working on the miRNA sequencing for an uncommon plant. I received data from a service company that gave to me a fastq file.
    I already did RNA-seq analysis, so I'm quite familiar with several tools such as FastQC, trimmomatic, Bowtie2, cuffdiff etc.

    I removed the 3' and 5' adapters, provided to me by the service company. The quality control confirmed that the adaptors sequences were right. I used cutadapt to remove the adapters. I have great peaks between 19 and 39 bp, also some reads between 39 and 51 (original reads length with adaptors attached).

    I downloaded the hairpin.fa file from MirBase, without filtering for a specific organism, changing all U in T and removing lines with strange chars (Y, K etc...).
    First strange thing:the alignment rate is very low, about 3%!

    So, I did the alignment again, this time versus the A. thaliana genome. The alignment rate increased to 20%.
    Second strange thing: if I launch htseq-count in order to count alignments, I found 0 for all mirnas!

    I'm sure that I'm wrong in some analysis steps...can someone help me?

    Thanks in advance
  • wynstep
    Member
    • Jan 2014
    • 11

    #2
    Originally posted by wynstep View Post
    Dear Colleagues,
    I want to share my strange experience with you, to ask your opinions and help.

    I'm working on the miRNA sequencing for an uncommon plant. I received data from a service company that gave to me a fastq file.
    I already did RNA-seq analysis, so I'm quite familiar with several tools such as FastQC, trimmomatic, Bowtie2, cuffdiff etc.

    I removed the 3' and 5' adapters, provided to me by the service company. The quality control confirmed that the adaptors sequences were right. I used cutadapt to remove the adapters. I have great peaks between 19 and 39 bp, also some reads between 39 and 51 (original reads length with adaptors attached).

    I downloaded the hairpin.fa file from MirBase, without filtering for a specific organism, changing all U in T and removing lines with strange chars (Y, K etc...).
    First strange thing:the alignment rate is very low, about 3%!

    So, I did the alignment again, this time versus the A. thaliana genome. The alignment rate increased to 20%.
    Second strange thing: if I launch htseq-count in order to count alignments, I found 0 for all mirnas!

    I'm sure that I'm wrong in some analysis steps...can someone help me?

    Thanks in advance
    Anyone helps me? Please!

    Comment

    • NextGenSeq
      Senior Member
      • Apr 2009
      • 482

      #3
      The Illumina miRNA library kit is known to display ligation bias. There is probably something wrong with your library.

      Comment

      • wynstep
        Member
        • Jan 2014
        • 11

        #4
        Originally posted by NextGenSeq View Post
        The Illumina miRNA library kit is known to display ligation bias. There is probably something wrong with your library.

        http://www.ncbi.nlm.nih.gov/pubmed/22647250
        Thank you very much for your help!
        So, what is your suggestion? How to proceed to remove or reduce ligation biases?

        Thank you!

        Comment

        • wynstep
          Member
          • Jan 2014
          • 11

          #5
          If someone wants, I can attach the fastqc files after 3' adaptor trimming...in order to have a better overview of my strange situation. I hope someone can help me, cause I finished the ideas on how to solve this problem.

          Tried the adaptor trimming with: trimmomatic, cutadapt, fasts_clipper, novoalign etc...
          Tried mapping with: bowtie, bowtie2, mirdeep2 etc...
          for now I only want to know if there are some known mirnas...

          The only thing I did not try is BLAST.

          Please help!

          Comment

          • NextGenSeq
            Senior Member
            • Apr 2009
            • 482

            #6
            Originally posted by wynstep View Post
            Thank you very much for your help!
            So, what is your suggestion? How to proceed to remove or reduce ligation biases?

            Thank you!
            The paper at the link describes how to reduce ligation bias.

            The Bioo Small RNA kit uses this method for Illumina platforms.

            Ion Torrent has used that method for a couple years for the PGM and Proton sequencers.

            Comment

            • Anton1
              Junior Member
              • Sep 2010
              • 1

              #7
              Seems quite normal to me since major population of sRNAs in Plants (like A thaliana) are not miRNAs but siRNAs (a mixture of 21, 22 and 24 mers) not well conserved and arranged along the genome in cluster. I guess that you have got the mir390 mir168 and others in your mapped miRNAs since they are well conserved in thaliana as well as particular cluster of siRNA, also conserved.

              Comment

              • wynstep
                Member
                • Jan 2014
                • 11

                #8
                Originally posted by NextGenSeq View Post
                The paper at the link describes how to reduce ligation bias.

                The Bioo Small RNA kit uses this method for Illumina platforms.

                Ion Torrent has used that method for a couple years for the PGM and Proton sequencers.
                I've read the paper you suggested, but I didn't find any bioinformatics suggestion on how to treat raw data from sequencing "affected" by Illumina adaptors ligation biases... Am I missing something important into the paper or are they focusing only on a sperimental solution (only on library preparation I mean)?

                Thanks for your help!

                Comment

                • kerplunk412
                  Senior Member
                  • Jun 2012
                  • 119

                  #9
                  Hi wynstep,
                  I have seen that low-mapping libraries can sometimes be attributed to some sort of artifact product that is taking up many of your reads. If this artifact is present in many of your reads, you should be able to find it with FastQC in the overrepresented sequences section. You will probably need to do this after adapter trimming, as otherwise I think the only overrepresented sequences that will be reported are from the 3' adapter. Also, you may want to just try BLASTing some random sequences from your data to see if you can get an indication of what they represent.

                  Comment

                  Latest Articles

                  Collapse

                  • SEQadmin2
                    Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                    by SEQadmin2



                    CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                    Despite this, “CRISPR helped turn genome editing from a specialized technique into
                    ...
                    Today, 11:01 AM
                  • SEQadmin2
                    Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                    by SEQadmin2


                    Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                    The systematic characterization of the human proteome has
                    ...
                    07-20-2026, 11:48 AM
                  • SEQadmin2
                    Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                    by SEQadmin2



                    Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                    ...
                    07-09-2026, 11:10 AM

                  ad_right_rmr

                  Collapse

                  News

                  Collapse

                  Topics Statistics Last Post
                  Started by SEQadmin2, Today, 02:55 AM
                  0 responses
                  8 views
                  0 reactions
                  Last Post SEQadmin2  
                  Started by SEQadmin2, 07-24-2026, 12:17 PM
                  0 responses
                  12 views
                  0 reactions
                  Last Post SEQadmin2  
                  Started by SEQadmin2, 07-23-2026, 11:41 AM
                  0 responses
                  12 views
                  0 reactions
                  Last Post SEQadmin2  
                  Started by SEQadmin2, 07-20-2026, 11:10 AM
                  0 responses
                  24 views
                  0 reactions
                  Last Post SEQadmin2  
                  Working...