Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • ndeshpan
    Member
    • Nov 2009
    • 29

    #1

    HiSeq small rna data adapter trimming using Adapter_trim.pl (mirTools)

    Hi,

    My HiSeq data for small RNA looks like this


    @HWI-ST705:254:C0G8HACXX:5:1101:1674:1995 1:N:0:AGTCAA
    TGAGATGAAGCACTGTAGCTCTGGAATTCTCGGGT
    +
    CCCFFFFFHHHHHJJJJJJJIJJJJJJJJJJJJJH
    @HWI-ST705:254:C0G8HACXX:5:1101:1765:1986 1:N:0:AGTCAA
    TGAGAACTGAATTCCATAGGCTGTTGGAATTCTCG
    +
    BCCFFFFDHFHHHIJJJIJJJJJHHJJEHIJIJJJ
    @HWI-ST705:254:C0G8HACXX:5:1101:1785:1990 1:N:0:AGTCAA
    GCTCTGTGATGAACCCTGGAATTCTCGGGTGCCAA
    +
    =?@DFFFA=AADFHIJJJJGAHHGIIIJI?CFGHC
    @HWI-ST705:254:C0G8HACXX:5:1101:1825:1999 1:N:0:AGTCAA
    TTTGGCAATGGTAGAACTCACACCTGGAATTCTCG
    +
    CCCFFFFFHHHHHJJJJJJJJJJJJJJJJJJJJJJ
    @HWI-ST705:254:C0G8HACXX:5:1101:2182:1985 1:N:0:AGTCAA
    CAACNGAATCCCAAAAGCAGCTGTGGAATTCTCGG
    +
    @@@D#2=BDDHFHBGHHIIIGHHHIGGGHH<?DHI
    @HWI-ST705:254:C0G8HACXX:5:1101:2106:1988 1:N:0:AGTCAA
    TAGCTTATCAGACTGATGTTGACTTGGAATTCTCG
    +
    ??@FFFD+=CFFFHGIJJGIHHHHHJCFHEHHHDH
    @HWI-ST705:254:C0G8HACXX:5:1101:2543:1995 1:N:0:AGTCAA
    TTCACAGTGGCTAAGTTCTGCTGGAATTCTCGGGT
    +
    CCCFFFFFHHHHHJJJJJJJJJJJJJJJJJJJJJH

    When I use the Adapter_trim.pl script from miRTools with format option "3" (for illumina format 1.3+) or evn "2" for the older formats ..I get a empty output file..

    The previous illumina datasets had the complete IDs repeated befire the quality value lines and the script used to work good for me...

    Any suggestions?

    regards,

    Nandan
    The read grouping script from miRAnalyser also gives me issues for HiSeq dataset (again this worked well for GAII illumina data)
  • maasha
    Senior Member
    • Apr 2009
    • 153

    #2
    I don't know miRTools, but I am pretty convinced that Biopieces will be quite helpful for this. Especially find_adaptor.

    Comment

    • vikas0633
      Junior Member
      • Mar 2012
      • 3

      #3
      To my knowledge - best way to do adapter filtering/trimming is



      if you are not fan of unix system then try galaxy
      Galaxy is a community-driven web-based analysis platform for life science research.

      Comment

      • sgcsd
        Junior Member
        • Aug 2012
        • 7

        #4
        Hi all:
        I am new to Illumina sequencing. I have a very basic question. The sequence we get after running through the Illumina pipeline, does they contain adapters for all the reads or only few reads.

        Recently we did an sequencing run through Hiseq2000 (multiplexed) and the fastq file has only few reads containing (5%) adapters or primers. I used the adapter and primer sequences used in library prep (from illumina truseq).

        I read some where that when the pipeline demultiplex it trims the reads and removes the barcode.Is it true.

        Please reply or direct me to some literature that explains the basic.

        Thank you

        Comment

        • bharat_iyengar
          Member
          • Dec 2012
          • 20

          #5
          i am facing a similar problem..

          I was interested in getting some information from a publicly available hiseq2000 small RNA seq data from drosophila.
          However, the library was prepared by cloning and not using truseq (as reported in the SRA. Accession number SRR513393).

          This isnt a great concern. I used fastqc to analyse the reads and the quality distribution seemed to be pretty okay (PFA). However, no overrepresented sequence was detected and I am unsure of the sequence of adapters. The reads are 50 nt long (more than twice the size of any miRNA or similar RNAs).

          I used bowtie v0.12.9 to align the reads against the drosophila transcriptome index that I built from flybase transcripts release v5.49, with options (-v 2 --norc -a --best --strata). No read got aligned, and I suspect that it might be because of some bogus sequence filling up the ends. I am not able to detect what those bogus sequences might be.

          Any tips for preprocessing.

          Plus, when I used tophat to align the reads against genome index with annotations provided from GFF file v5.49 from flybase, then tophat stopped with a report "gtf_to_fasta returned an error" [ isn't tophat supposed to accept GFF v3 files ??]
          Attached Files

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
            by SEQadmin2



            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

            Despite this, “CRISPR helped turn genome editing from a specialized technique into
            ...
            Yesterday, 11:01 AM
          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, Yesterday, 02:55 AM
          0 responses
          8 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-24-2026, 12:17 PM
          0 responses
          12 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-23-2026, 11:41 AM
          0 responses
          12 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-20-2026, 11:10 AM
          0 responses
          24 views
          0 reactions
          Last Post SEQadmin2  
          Working...