Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • starbug16
    Junior Member
    • Feb 2014
    • 8

    #1

    Over-represented sequences

    Hi,

    I have one lane (~360 million reads) of Illumina transcriptome data which I am trying to assemble. Having only a tiny server, I am attempting to reduce this dataset with digital normalization. However, before doing this, I have run fastqc which flags up some over-represented sequences. When I BLAST these, I get matches to my organism's 16S, 18S and 28S rRNA genes.

    As a bit of a beginner, I was wondering whether this was normal and whether I need to worry or do anything about this?

    Any help or advice very much appreciated!!
  • dpryan
    Devon Ryan
    • Jul 2011
    • 3478

    #2
    That's completely normal, don't give it a second thought.

    Comment

    • mastal
      Senior Member
      • Mar 2009
      • 666

      #3
      It's quite normal, especially if the library prep was done using total RNA, and not using any methods to reduce or remove ribosomal RNA.

      Comment

      • starbug16
        Junior Member
        • Feb 2014
        • 8

        #4
        Phew, thank you for the responses. Mind much at rest and now onwards to diginorm-ing!

        Comment

        • yueluo
          Member
          • Aug 2013
          • 82

          #5
          If you're not interested in ribosomal genes, you might as well remove all those reads.

          Comment

          • Marianna85
            Member
            • Mar 2012
            • 32

            #6
            Hi everybody,

            actually I had the same problem (that is not really a problem but something that is quite normal) and I want to ask you an opinion that it might be helpful also for starbug16.

            I did several libraries by using different kits and I had different percentages of reads mapping to rRNAs, ranging from 1 to 30%.
            I think that the problem arises when you have to compare samples having very different percentages of rRNAs. I mean: if I have to find DE between 2 samples having respectively 15% and 30% of rRNA reads, I have the impression that final results are biased by the very different percentages of rRNAs (which compete for the sequencing and affect the number of mRNA sequences). Did you understand my point?
            In this case, do you think that a normalization procedure (like TMM or DESeq) will minimize this issue?

            Thank you to anybody who will tell his opinion!

            Marianna

            Comment

            • yueluo
              Member
              • Aug 2013
              • 82

              #7
              I usually remove reads that map to rRNA(and/or other sources of contamination) , then proceed with mapping/DE-analysis .

              Comment

              • Marianna85
                Member
                • Mar 2012
                • 32

                #8
                Hi Yueluo,
                so you think it is not a problem if your samples have really different percentages of rRNAs??

                Marianna

                Comment

                Latest Articles

                Collapse

                • SEQadmin2
                  New Genomics Technologies Take Aim at Long-Standing Limits
                  by SEQadmin2


                  Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

                  We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
                  ...
                  Yesterday, 10:25 AM
                • SEQadmin2
                  How Immunogenomics Decodes Immunity’s Genetic Blueprint
                  by SEQadmin2




                  The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

                  This convergence of genetics, immunology, and computation...
                  09-01-2026, 05:41 AM

                ad_right_rmr

                Collapse

                News

                Collapse

                Topics Statistics Last Post
                Started by SEQadmin2, Today, 09:51 AM
                0 responses
                7 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 09-25-2026, 09:06 AM
                0 responses
                31 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 09-23-2026, 11:05 AM
                0 responses
                26 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 09-18-2026, 11:37 AM
                1 response
                47 views
                0 reactions
                Last Post pekgio
                by pekgio
                 
                Working...