Hello, I am new to analysing RNAseq data and I'm learning how to run bioinformatics just by self-reading.
The samples were paired-end reads, sequenced by Illumina HISEQ 2000 and then I ran the FastQ files via FastQC and I generated the following report:
Sequence tenth: 101
I received a "RED CROSS" for per base sequence content and I've attached the graph for your information.
I am so received a "RED CROSS" for sequence duplication levels.
The over-represented sequences are:
GTGGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGT
GGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGTTA
GCGAATGGCTCATTAAATCAGTTATGGTTCCTTTGGTCGCTCGCTCCTCT
However, I have no idea where they're coming from because they do not match my adapter sequences (GGAGAA) and also FasQC does not report them as 'adapter sequences'.
I'll show the K'mer content graph as well, with most over-represented kMERS at the 3' end.
I wonder what could be the source of these over-represented sequences and how best to deal with them?
Thank you very much for your advice.
The samples were paired-end reads, sequenced by Illumina HISEQ 2000 and then I ran the FastQ files via FastQC and I generated the following report:
Sequence tenth: 101
I received a "RED CROSS" for per base sequence content and I've attached the graph for your information.
I am so received a "RED CROSS" for sequence duplication levels.
The over-represented sequences are:
GTGGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGT
GGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTGTCTGGTTA
GCGAATGGCTCATTAAATCAGTTATGGTTCCTTTGGTCGCTCGCTCCTCT
However, I have no idea where they're coming from because they do not match my adapter sequences (GGAGAA) and also FasQC does not report them as 'adapter sequences'.
I'll show the K'mer content graph as well, with most over-represented kMERS at the 3' end.
I wonder what could be the source of these over-represented sequences and how best to deal with them?
Thank you very much for your advice.
Comment