Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • windnature03
    Junior Member
    • Nov 2014
    • 3

    Strange per base sequence content from FastQC report

    Hi all ,

    I am new to bioinformatics. I have encountered some problems with the data analysis(RNA-seq for human, Pair-end 101cycle). The per base sequence content and Kmer content look pretty strange from what i expected.
    Here are my questions (please use the attached file for reference)

    1.There is a sudden rise of %A around 50bp. Would it be adapter contamination? but the adapter content keeps low throughout the whole run.

    2. What is the possible cause of the A/T imbalance?

    3. What is the possible cause of peaks around 40-49bp from the Kmer content?

    4. why the base quality drops after 50bp?

    Can anyone give me some clue on these questions, it's been puzzling me for a week.

    Thank you
    Attached Files
  • wdecoster
    Member
    • Oct 2015
    • 97

    #2
    For question 1, the increase in %A is not too bad. I think it can be explained that for some reads you start running in the polyA tail, which you might want to trim off.

    Comment

    • nucacidhunter
      Jafar Jabbari
      • Jan 2013
      • 1250

      #3
      Would you know what kit was used for library prep. Some kits adapter are different from the ones that FastQC can detect.

      Comment

      • windnature03
        Junior Member
        • Nov 2014
        • 3

        #4
        Hi nucacidhunter and wdecoster,

        TruSeq RNA Library Prep Kit v2 was used, the link below shows the overpresented adapters sequence

        Discover the magic of the internet at Imgur, a community powered entertainment destination. Lift your spirits with funny jokes, trending memes, entertaining gifs, inspiring stories, viral videos, and so much more from users.



        This is the result from bioanalyzer, the peak lies around 250-300bp. Subtracting the length of the adapter (60bp), the insert should be around 120-130bp. In my opinion, it is less likely for a adatper sequence to be read at 50 cycle .

        Discover the magic of the internet at Imgur, a community powered entertainment destination. Lift your spirits with funny jokes, trending memes, entertaining gifs, inspiring stories, viral videos, and so much more from users.



        Thank you
        Last edited by windnature03; 03-01-2017, 01:51 AM. Reason: correcting image link

        Comment

        • nucacidhunter
          Jafar Jabbari
          • Jan 2013
          • 1250

          #5
          Originally posted by windnature03 View Post
          Hi all ,

          I am new to bioinformatics. I have encountered some problems with the data analysis(RNA-seq for human, Pair-end 101cycle). The per base sequence content and Kmer content look pretty strange from what i expected.
          Here are my questions (please use the attached file for reference)

          1.There is a sudden rise of %A around 50bp. Would it be adapter contamination? but the adapter content keeps low throughout the whole run.

          2. What is the possible cause of the A/T imbalance?

          3. What is the possible cause of peaks around 40-49bp from the Kmer content?

          4. why the base quality drops after 50bp?
          1 and 2- One possible explanation is 3' bias due to input RNA low quality which has increased polyA representation.

          3- Sequences TATGCCG and CGTATGC are over-represented Kmers with TATGC overlap. You might check to see if they are from a particular highly expressed gene or spike in RNA if it was used.

          4- It does not seem to be library related. You can ask the sequencing centre for an explanation. They can look at other lanes in the same flow cell to see if sequencing reagent or sequencer had any issues.

          Comment

          • GenoMax
            Senior Member
            • Feb 2008
            • 7142

            #6
            Originally posted by windnature03 View Post

            4. why the base quality drops after 50bp?
            It is possible that inserts in your library are smaller than what you had expected. This generally causes adapter read-through and results in Q-score drops.

            Have you scanned/trimmed this data for presence of adapters? I recommend you try bbduk.sh from BBMap suite for that purpose. There are threads on SeqAnswers that will guide you on how to use bbduk. You can also use bbmerge.sh or bbmap.sh (if you have a reference genome) from the same suite to estimate your library insert size.

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              Yesterday, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM
            • SEQadmin2
              Cancer Drug Resistance: The Lingering Barrier to Rising Survival
              by SEQadmin2



              Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

              There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
              07-08-2026, 05:17 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Yesterday, 11:10 AM
            0 responses
            9 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-13-2026, 10:26 AM
            0 responses
            30 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-09-2026, 10:04 AM
            0 responses
            40 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-08-2026, 10:08 AM
            0 responses
            25 views
            0 reactions
            Last Post SEQadmin2  
            Working...