Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • ErikFas
    Member
    • Jun 2014
    • 86

    Alignment of small RNA data

    I was recently at a meeting about RNA-seq in general, and the topic of small RNA-seq came up, something with which I'm quite unfamiliar. The discussions were interesting, but seeing as I didn't know much about sRNA-seq (and I was the "RNA-seq"-guy at the meeting), they didn't get very far. I've since tried to learn a bit about it, and I wanted to ask some questions to clear up things I'm not sure about...

    1) A general pipeline for sRNA-seq. As far as I understand it, the sequencing adapters are proportionally a much larger part of the reads than for normal RNA-seq. This would make adapter trimming more or less mandatory for any sRNA-seq analysis. Is this correct?

    2) Seeing as sRNA is a lot smaller, would that mean that there are more duplicated reads in an sRNA-seq dataset? If so, would you remove them?

    3) As far as alignment goes, I can't really understand if one should use one of the sRNA-specific aligners I seem to find by googling, or to use one of the normal RNA-seq aligners (STAR, Tophat, etc.). I seem to find information saying that you can use either...

    4) Can you align to the normal human reference genome (such as GRCh38), or do you need to add some sRNA-specific database? I found miRBase, for example, which (as far as I can tell) is a database for miRNA sequences. I assume one could align to that, if one is only interested in miRNA? Or should those sequences be added to e.g. GRCh38 and then aligned to the collated reference?

    Since I'm interested in this purely from a learning and knowledge perspective, I won't actually work with any sRNA-seq dataset. I did download a run from the SRA and put it through my standard alignment pipeline just to see what happened, though. I got around 80% ambigously alignments and about 10% duplicated reads using just a very simple STAR 2-pass alignment to GRCh38 without any sRNA-specific sequences added and no adapter/quality trimming. Do these numbers make sense for the non-optimised (from an sRNA perspective) pipeline used? What would be required to get a better alignment?
  • nanos
    Member
    • May 2010
    • 11

    #2
    Dear Eric, we are very often analyzing sRNA data and I can give you some insight.
    1) Adaptor trimming is really a must. With the minimum sequencing length being 50 you always have adaptor remnants in the sequence.

    2) removing duplicated reads would be a problem. The problem is, that you will in most of the cases have the full length sequence of you the sRNA sequenced. Therefore in contrast to RNAseq you do not have a random shifting in you sequence (hope this is understandable). Removing duplicates will leave you most likely with a very low and very similar count for all the miRNAs no matter how high/different they were expressed. You can use adaptors containing random nucleotides and then use these 8Ns in combination with the sRNA sequence to assess the duplication rate.

    3) we use good old bowtie and it works perfectly fine for us (if there are different opinions on that one, any input is appreciated)


    4) I guess the answer here largely depends on your question.

    hope that helps as a start.

    Comment

    • lre1234
      Senior Member
      • Aug 2011
      • 110

      #3
      Hi,
      I do a lot of short RNA-seq and here are some thoughts (but there are other ways of doing things that work well):

      1. Agreed that adapter trimming is a must or most of your reads will not map. We use cutadapt which works really nice.
      2. No duplicate read removing is needed nor should be done. You'll loose lots of things.

      3. bowtie works well, I have also used BWA which also seemed to work well but usually default to bowtie. As far as I understand, STAR wouldn't work for short RNAs as it was designed for long RNA and specifically paired-end (but don't quote me here, I may be wrong). STAR is our goto aligner for long RNA.

      4. As far as aliging. In my opinion, you should always align to the whole genome (GRCh37 or 38, which ever you choose). Afterwords intersect with miRBase or some other database of interest. Also, keep in mind, that the vast majority of miRNAs are 'unique' sequences in genome and should align uniquely. But there are cases, in which some miRNAs have duplicate sequences in the genome (e.g. miR-92a-3p, or miR-1302 which the same sequence is in 11 places in the genome). Also by mapping to the whole genome, you could do additional things like novel miR discovery. Some people do use miRBase sequences and align to those instead of the whole genome, but I personally think that is a bad idea, and will give a false-sense of what you are looking at. Essentially, you would be 'forcing' many reads to align to those regions, when in fact they would align better to other places in the genome, especially when you allow a mismatch in there.

      Have fun with it. miRNAs do lots of interesting things and have many useful roles!

      Comment

      • manwar
        Junior Member
        • Nov 2017
        • 1

        #4
        Which GTFs to use for annotation of sRNA?

        Hello everyone,

        Following on from ErikFas's query about using the normal human reference genome for sRNA-seq analysis, I wanted to ask if a regular gtf/gff from Ensembl or UCSC can be used for annotation purposes of sRNA or are there specific gtfs?

        Thanks a lot!

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
          by SEQadmin2


          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

          The systematic characterization of the human proteome has
          ...
          07-20-2026, 11:48 AM
        • SEQadmin2
          Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
          by SEQadmin2



          Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
          ...
          07-09-2026, 11:10 AM
        • SEQadmin2
          Cancer Drug Resistance: The Lingering Barrier to Rising Survival
          by SEQadmin2



          Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

          There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
          07-08-2026, 05:17 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, Today, 12:17 PM
        0 responses
        10 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, Yesterday, 11:41 AM
        0 responses
        11 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-20-2026, 11:10 AM
        0 responses
        23 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-13-2026, 10:26 AM
        0 responses
        37 views
        0 reactions
        Last Post SEQadmin2  
        Working...