Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • bbm
    Member
    • Sep 2011
    • 38

    #1

    how to increase the mapping rate?

    I have a set of RNA-seq dataset of single end 100bp reads (30 million per sample), and first using tophat2, mapping rate is only 5% to the ref genome. Then I tried to trim raw data to 40-100bp, and mapping rate increase to 18%. I'm doing the mapping with no trimmed data right now...

    I wonder what other ways I can try to increase the mapping rate? trim read range to 50-100? increase the phred score based on fastqc?

    Any comments will be appreciated!
  • GenoMax
    Senior Member
    • Feb 2008
    • 7142

    #2
    Can you post the FastQC plots of what the data looks like? No point in doing random trimming of data.

    Take a few reads and do an old fashioned blast to make sure the data is from your sample/correct genome. Mistakes sometimes happen at sequencing cores.

    Comment

    • bbm
      Member
      • Sep 2011
      • 38

      #3
      I just had no trimming data alignment, and it is 15%.

      15.66% overall alignment rate

      I will post the fastqc plots soon. Thank you!

      Comment

      • bbm
        Member
        • Sep 2011
        • 38

        #4
        Attached here is the fastqc before I trimmed
        Attached Files

        Comment

        • bbm
          Member
          • Sep 2011
          • 38

          #5
          This is the fastqc after I trimmed using trimmomatic, 40-100bp

          java -jar /usr/local/apps/trimmomatic/Trimmomatic-0.32/trimmomatic-0.32.jar SE 1.fastq 1.trimmed.fastq ILLUMINACLIP:/usr/local/apps/trimmomatic/Trimmomatic-0.32/adapters/TruSeq3-SE.fa:2:30:10 LEADING:3 TRAILING:3 SLIDINGWINDOW:4:15 MINLEN:40
          Attached Files

          Comment

          • GenoMax
            Senior Member
            • Feb 2008
            • 7142

            #6
            Q-score wise there is no issue, so the problem must lie elsewhere. It is possible to get great data that may not align at all so this is only part of the QC. Report back on the blast result. Do the GC plots look strange?

            Comment

            • Brian Bushnell
              Super Moderator
              • Jan 2014
              • 2709

              #7
              I don't really understand what you mean by "trimming to 40-100bp". But, it would not surprise me if your problem was adapter contamination; do you know what kind of adapters were used? They might not be TruSeq.

              Comment

              • WhatsOEver
                Senior Member
                • Apr 2012
                • 215

                #8
                What organism are you working with and what is your reference?
                I have seen such fastqc results just recently.
                The reason was a severe rRNA contamination. Maybe mRNA enrichment / ribo-depletion didn't work (or wasn't done)?
                If the respective sequences are not (or are only partially) represented in your reference, you can of course not map to them. Look at the sequence duplication levels: if there is an increase at 10k, this is an indication for that. If you are working with human samples, the relatively high GC content is another one.
                To verify this, simply use the rRNA sequences as reference and map to them.

                Comment

                • bbm
                  Member
                  • Sep 2011
                  • 38

                  #9
                  Originally posted by GenoMax View Post
                  Q-score wise there is no issue, so the problem must lie elsewhere. It is possible to get great data that may not align at all so this is only part of the QC. Report back on the blast result. Do the GC plots look strange?
                  here is the overall fastqc
                  Attached Files

                  Comment

                  • bbm
                    Member
                    • Sep 2011
                    • 38

                    #10
                    Originally posted by WhatsOEver View Post
                    What organism are you working with and what is your reference?
                    I have seen such fastqc results just recently.
                    The reason was a severe rRNA contamination. Maybe mRNA enrichment / ribo-depletion didn't work (or wasn't done)?
                    If the respective sequences are not (or are only partially) represented in your reference, you can of course not map to them. Look at the sequence duplication levels: if there is an increase at 10k, this is an indication for that. If you are working with human samples, the relatively high GC content is another one.
                    To verify this, simply use the rRNA sequences as reference and map to them.
                    The reference is honeybee genome, which is the 2nd version so far. Thank you for your suggestion. I think it may be the problem of low quality lib prep.

                    Comment

                    • bbm
                      Member
                      • Sep 2011
                      • 38

                      #11
                      Originally posted by Brian Bushnell View Post
                      I don't really understand what you mean by "trimming to 40-100bp". But, it would not surprise me if your problem was adapter contamination; do you know what kind of adapters were used? They might not be TruSeq.
                      The lib was done by the NEBNext® RNA Library Prep Kit for Illumina, so it should be TruSeq adaptors.

                      Comment

                      • WhatsOEver
                        Senior Member
                        • Apr 2012
                        • 215

                        #12
                        From looking at the fastqc output (btw: there is a new, slightly better fastqc version available), I can only say again that it looks very similar to our rRNA "contaminated" samples. More interestingly, we also used the NEB kit...

                        Comment

                        • GenoMax
                          Senior Member
                          • Feb 2008
                          • 7142

                          #13
                          @bbm: Were you mapping to the entire genome or just the transcriptome?

                          Comment

                          Latest Articles

                          Collapse

                          • SEQadmin2
                            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                            by SEQadmin2



                            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                            Despite this, “CRISPR helped turn genome editing from a specialized technique into
                            ...
                            Today, 11:01 AM
                          • SEQadmin2
                            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                            by SEQadmin2


                            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                            The systematic characterization of the human proteome has
                            ...
                            07-20-2026, 11:48 AM
                          • SEQadmin2
                            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                            by SEQadmin2



                            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                            ...
                            07-09-2026, 11:10 AM

                          ad_right_rmr

                          Collapse

                          News

                          Collapse

                          Topics Statistics Last Post
                          Started by SEQadmin2, Today, 02:55 AM
                          0 responses
                          6 views
                          0 reactions
                          Last Post SEQadmin2  
                          Started by SEQadmin2, 07-24-2026, 12:17 PM
                          0 responses
                          11 views
                          0 reactions
                          Last Post SEQadmin2  
                          Started by SEQadmin2, 07-23-2026, 11:41 AM
                          0 responses
                          12 views
                          0 reactions
                          Last Post SEQadmin2  
                          Started by SEQadmin2, 07-20-2026, 11:10 AM
                          0 responses
                          24 views
                          0 reactions
                          Last Post SEQadmin2  
                          Working...