Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • aurorasea1
    Junior Member
    • May 2016
    • 5

    #1

    Strange looking fastQC per base sequence content, duplications, overrepresented, KMER

    Hello, I am new to analysing RNAseq data and I'm learning how to run bioinformatics just by self-reading.
    The samples were paired-end reads, sequenced by Illumina HISEQ 2000 and then I ran the FastQ files via FastQC and I generated the following report:

    Sequence tenth: 101
    I received a "RED CROSS" for per base sequence content and I've attached the graph for your information.

    I am so received a "RED CROSS" for sequence duplication levels.

    The over-represented sequences are:
    GTGGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGT
    GGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGTTA
    GCGAATGGCTCATTAAATCAGTTATGGTTCCTTTGGTCGCTCGCTCCTCT
    However, I have no idea where they're coming from because they do not match my adapter sequences (GGAGAA) and also FasQC does not report them as 'adapter sequences'.

    I'll show the K'mer content graph as well, with most over-represented kMERS at the 3' end.

    I wonder what could be the source of these over-represented sequences and how best to deal with them?
    Thank you very much for your advice.
    Attached Files
  • nucacidhunter
    Jafar Jabbari
    • Jan 2013
    • 1250

    #2
    They are 100% match to the following:
    GTGGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGT
    PREDICTED: Sinocyclocheilus rhinocerous neuroligin-3-like (LOC107733424), transcript variant X14, misc_RNA

    GGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGTTA
    PREDICTED: Capsicum annuum uncharacterized LOC107852466 (LOC107852466), ncRNA

    GCGAATGGCTCATTAAATCAGTTATGGTTCCTTTGGTCGCTCGCTCCTCT
    Uncultured eukaryote 18S ribosomal RNA, partial sequence

    Red crosses are intended to indicate shotgun library quality and most of the time they will show up in RNA-Seq data. Duplication rate is high but that would be dependent on the quantity of input RNA.

    Comment

    • aurorasea1
      Junior Member
      • May 2016
      • 5

      #3
      Originally posted by nucacidhunter View Post
      They are 100% match to the following:
      GTGGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGT
      PREDICTED: Sinocyclocheilus rhinocerous neuroligin-3-like (LOC107733424), transcript variant X14, misc_RNA

      GGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGTTA
      PREDICTED: Capsicum annuum uncharacterized LOC107852466 (LOC107852466), ncRNA

      GCGAATGGCTCATTAAATCAGTTATGGTTCCTTTGGTCGCTCGCTCCTCT
      Uncultured eukaryote 18S ribosomal RNA, partial sequence

      Red crosses are intended to indicate shotgun library quality and most of the time they will show up in RNA-Seq data. Duplication rate is high but that would be dependent on the quantity of input RNA.
      Thanks very much but I'm very surprised how the first 2 ended up in my sequence (!)
      But would they cause problem in my subsequent analysis? Because I'll be mapping them to the human genome rather than de novo assembly, I suppose that'd exclude these 'non-mappable' reads?
      Thank you

      Comment

      • Michael.Ante
        Senior Member
        • Oct 2011
        • 127

        #4
        Hi,
        it looks, as you have a random-primer in the first 8-9 bases. It might be useful to either trim this sequence or increase allowed mismatches in your alignment.

        Regarding the over-represented sequences, they might or might not map to your target genome. Nevertheless, if you have low alignment rates, you may want to check the unmapped reads. Some aligners like TopHat2 store these automatically, others like the STAR aligner need a parameter for saving them.

        Cheers,

        Michae

        Comment

        • GenoMax
          Senior Member
          • Feb 2008
          • 7142

          #5
          Best solution may be to do nothing (except scan for presence of adapters and remove them if necessary).

          Take a look at these two posts. RNAseq bias and Over-representation.

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
            by SEQadmin2



            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

            Despite this, “CRISPR helped turn genome editing from a specialized technique into
            ...
            07-31-2026, 11:01 AM
          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, 08-03-2026, 10:13 AM
          0 responses
          15 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-31-2026, 02:55 AM
          0 responses
          32 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-24-2026, 12:17 PM
          0 responses
          23 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-23-2026, 11:41 AM
          0 responses
          21 views
          0 reactions
          Last Post SEQadmin2  
          Working...