Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • buthercup_ch
    Member
    • Apr 2014
    • 41

    #1

    Low reads mapping in RNA-Seq analysis

    Dear all.
    After starting to work a little while ago with RNA in prokaryotes in order to perform RNA-Seq analysis, I finally arrived to the moment of the bioinformatic analysis of the obtained reads.
    I prepared the cDNA library by following a modified TruSeq protocol for Illumina and the quality of the preparation by analysis using DNA High Sensitivity chip in Bioanalyzer was very good.
    After sequencing reaction, I used CLC Genomics Workbench to perform the RNA-Seq analysis. First of all, I run the tool for checking the quality of the reads, and the sequencing reaction seemed to be almost perfect, so I was very happy. But when running the RNA-Seq Analysis tool included in the software (using default parameters) it happens that more than 60% of the reads doesn't match with my reference genome.
    I have to say that as reference genome I use a multifasta file containing the list of all CDS, but not the assemble annotated genome.
    I was wondering why just 30% of the reads are mapped during the analysis…. Now I think that it might be due to the use of a CDS list… that would make all reads falling in between two CDS or intergenic regions will not be mapped. Am I right? Have any of you any other suggestion?
    Thank you very much in advance!!
  • GenoMax
    Senior Member
    • Feb 2008
    • 7142

    #2
    How long are your "reference" fasta sequences? Since this is a prokaryote you do not need to account for introns so in theory the alignments should be simpler. Have you tried use the "align to reference" workflow instead of RNA-seq under transcriptome analysis in CLC?

    Does FastQC (http://www.bioinformatics.babraham.a...ojects/fastqc/) give the data a reasonably clean bill of health? No over-represented primer dimers/adapters.

    Comment

    • WhatsOEver
      Senior Member
      • Apr 2012
      • 215

      #3
      @GenoMax: He writes "intergenic regions", not "introns".

      @buthercup_ch: Why do you use CDS sequences instead of a genome sequence (As you indicate that there is one available)? If the genome is poorly annotated, it is possible - although highly unlikely - that your reads map to yet unknown/not annotated genes. Instead, I would rather look for non-coding RNAs in your reads. But a mapping to the genome will tell you more.

      Comment

      • Brian Bushnell
        Super Moderator
        • Jan 2014
        • 2709

        #4
        There's a lot of weird junk in prok RNA-seq that does not map well, though normally it's under 10% of the reads. Still, mapping to the full genome, as suggested, with a splice-capable aligner will give a much better picture of what's happening. Introns are rare in prokaryotes, but self-splicing genes do exist. Furthermore, many bacteria are capable of modifying their own DNA to combat viruses (and alternately, can have their DNA modified by viruses) which creates reads that appear to have structural variations.

        It's also possible that you have some kind of contamination. I suggest BLASTing a thousand or so reads to NT and RefSeq Microbial to see what they are. They could be human, phiX, or some common bacterial contaminant, for example.

        Comment

        • swbarnes2
          Senior Member
          • May 2008
          • 910

          #5
          Well, what have you looked at? Have you pulled out some of the unmapped reads to examine them? Check their quality? BLASTed them? Done de novo assembly on them, to see if you can make a contig?

          Comment

          • alexdobin
            Senior Member
            • Feb 2009
            • 161

            #6
            I would bet this is caused by ribosomal RNA, which comprise the largest fraction of all reads, and do not ribo-deplete well with standard kits like Ribo-Zero (at least this is what happened to us a couple of years ago). Of course, they will not map to CDSs - mapping to the full genome will reveal them as multi-mappers.

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
              by SEQadmin2



              CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

              Despite this, “CRISPR helped turn genome editing from a specialized technique into
              ...
              07-31-2026, 11:01 AM
            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, 07-31-2026, 02:55 AM
            0 responses
            19 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-24-2026, 12:17 PM
            0 responses
            16 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-23-2026, 11:41 AM
            0 responses
            16 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-20-2026, 11:10 AM
            0 responses
            26 views
            0 reactions
            Last Post SEQadmin2  
            Working...