Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts

  • mmartin
    replied
    Originally posted by sp144 View Post
    Marcel, is version 1.5 up yet? I can only find 1.4.2 as the latest version. If not, when do you anticipate 1.5 to be out?
    Thanks!
    Taking your question as a motivation: I've just released cutadapt 1.5! As always, see https://code.google.com/p/cutadapt/ for the changelog and download it from PyPI. Or, even better, just use "pip install cutadapt". Here is a copy of the changelog:
    • Adapter sequences can now be read from a FASTA file. For example, write -a file:adapters.fasta to read 3' adapters from adapters.fasta. This works also for -b and -g. This fixes the long-standing issue #33. Note that cutadapt isn't really optimized for trimming dozens or even hundreds of adapters!
    • There is now an option --mask-adapter, which can be used to not remove adapters, but to instead mask them with N characters. Thanks to Vittorio Zamboni for contributing this feature!
    • U characters in the adapter sequence are automatically converted to T.
    • Add the option -u/--cut, which can be used to unconditionally remove a number of bases from the beginning or end of each read.
    • When the new option --quiet is used, no report is printed after all reads have been processed.
    • When processing paired-end reads, cutadapt now checks whether the reads are properly paired.
    • To handle paired-end reads, an option --untrimmed-paired-output was added.

    Leave a comment:


  • sp144
    replied
    You can do that in post-processing. Just put everything on one line using sed:

    sed 'N;N;N;s/\\n/\\t/g'

    then remove lines containing \t+\t and after change all \t to \n.

    Marcel, is version 1.5 up yet? I can only find 1.4.2 as the latest version. If not, when do you anticipate 1.5 to be out?
    Thanks!
    Last edited by sp144; 07-31-2014, 03:37 PM.

    Leave a comment:


  • foivos
    replied
    Here is what I get
    @BS-DSFCONTROL03:317:C3PGTACXX:2:1101:4054:2147 1:N:0:GCCAAT
    TTAGGAAGAGGATAACAATTNGAAACAGTTGCTAAAACTCTATATGC
    +
    CCCFFFFFGHHHHJJJJJJJ#4AHGGIJIJJIJIJJJJJJJJJJJJJ
    @BS-DSFCONTROL03:317:C3PGTACXX:2:1101:4107:2164 1:N:0:GCCAAT
    AGTACCCCATGGAC
    +
    ?1?DD?BDA:C;22
    @BS-DSFCONTROL03:317:C3PGTACXX:2:1101:4138:2178 1:N:0:GCCAAT
    ATCGACACTTCGAACGCACTTGCGGCCCCGGGTTCCTCCCGGGGCTACGCC
    +
    CCCFFFFFHHHHHJJJJJJJJJIJJGGJJ:FG-5@D>EEH<?A@/'5<;;B
    @BS-DSFCONTROL03:317:C3PGTACXX:2:1101:4219:2179 1:N:0:GCCAAT

    +

    @BS-DSFCONTROL03:317:C3PGTACXX:2:1101:4242:2199 1:Y:0:GCCAAT
    CATACAGGACTCTTTCGAGGCCCTC
    +
    ==>A+2@<+?+?22<A+23)@C+1=
    It keeps the identifier and the "+" and removes the adapter and the sequence.

    I want it to remove everyting and not leave any gaps...

    Leave a comment:


  • foivos
    replied
    No I will not remove the lines as described in stackoverflow.

    I will isolate the problem, as it is part of a pipeline and I will make a new post soon.

    Leave a comment:


  • mmartin
    replied
    Originally posted by foivos View Post
    I am using cutadapt 1.3 (I will update to the new version soon), but there is an issue. After trimming the adapter it leaves some empty lines.

    Is there a way not to leave empty lines?
    Are you talking about reads that have a length of zero? This will appear as empty lines in the output file. Use cutadapt's --minimum-length option and set it to 1 or some higher value to avoid getting empty reads.

    Do not do what is described in the stackoverflow link because it will break your FASTQ file.

    Leave a comment:


  • GenoMax
    replied
    Originally posted by foivos View Post
    Hi guys

    I am using cutadapt 1.3 (I will update to the new version soon), but there is an issue. After trimming the adapter it leaves some empty lines.

    Is there a way not to leave empty lines? I don't want to write a script that parses again the file and fixes it.

    Thank you in advance
    Following only addresses issue of removing empty lines (I assume the results file is otherwise ok). It may be safer to write to a temp file instead of overwriting the original: http://stackoverflow.com/questions/1...om-a-unix-file
    Last edited by GenoMax; 04-23-2014, 05:04 AM.

    Leave a comment:


  • foivos
    replied
    Hi guys

    I am using cutadapt 1.3 (I will update to the new version soon), but there is an issue. After trimming the adapter it leaves some empty lines.

    Is there a way not to leave empty lines? I don't want to write a script that parses again the file and fixes it.

    Thank you in advance
    Last edited by foivos; 04-23-2014, 04:43 AM.

    Leave a comment:


  • mmartin
    replied
    Originally posted by vishwesh View Post
    Hi,
    cutadapt -a TGGAATTCTCGGGTGCCAAGG -g GUUCAGAGUUCUACAGUCCGACGAUC input.fastq > output.fastq
    Cutadapt removes only one adapter per read, so you need to run it twice with each adapter or specify the option --times=2. Also, you should use specify the 3' adapter starting with a "^" like so: -g ^GTTCAGAG...

    And I found that the 5' adapter has U instead of T. Will that be fine?
    No, not in cutadapt versions up to 1.4.2. But since it's a very good idea to support this, I just added this feature to cutadapt: Starting with cutadapt 1.5, all Us will be automatically replaced with Ts in the adapter sequence.

    Leave a comment:


  • vishwesh
    replied
    Hi,
    I am using cutadapt for removing the adapter sequence. I have 2 adapter sequence.

    RNA 5Adapter (RA5)
    5 GUUCAGAGUUCUACAGUCCGACGAUC
    RNA 3?Adapter (RA3)
    5 TGGAATTCTCGGGTGCCAAGG

    The 1st one is 5' adapter and 2nd is 3' adapter.

    I am using the following command line to remove the adapter seq.

    cutadapt -a TGGAATTCTCGGGTGCCAAGG -g GUUCAGAGUUCUACAGUCCGACGAUC input.fastq > output.fastq

    Length Distribution I get
    Mean sequence length: 32.49 ± 10.53 bp
    Minimum length: 16 bp
    Maximum length: 51 bp
    Length range: 36 bp
    Mode length: 51 bp with 2,852,626 sequences


    And I found that the 5' adapter has U instead of T. Will that be fine?

    I tried replacing U with T GUUCAGAGUUCUACAGUCCGACGAUC > GTTCAGAGTTCTACAGTCCGACGATC and tried removing adapter sequence.

    cutadapt -a TGGAATTCTCGGGTGCCAAGG -g GTTCAGAGTTCTACAGTCCGACGATC input.fastq > output.fastq

    Length Distribution I get
    Mean sequence length: 31.26 ± 11.29 bp
    Minimum length: 1 bp
    Maximum length: 51 bp
    Length range: 51 bp
    Mode length: 51 bp with 2,805,271 sequences

    I get varied length distribution in both the cases. Which one should I choose..
    First is the command that I am using is right??

    Kindly let me know.

    Thanks in advance.

    Regards
    Vishwesh

    Leave a comment:


  • mmartin
    replied
    There is a section in the README about this:


    In short, just run:

    Code:
    cutadapt -a AGATCGGAAGAGC -o trimmed.1.fastq.gz reads.1.fastq.gz
    cutadapt -a AGATCGGAAGAGC -o trimmed.2.fastq.gz reads.2.fastq.gz
    See the other sections in the README if you need to do more specialized things.

    Leave a comment:


  • GenoMax
    replied
    Use "trim galore" which is a wrapper for cutadapt to simplify things.

    Leave a comment:


  • arcolombo698
    replied
    Cutadapt

    Hello.

    Yes this website was helpful.

    Does cutadapt take a Fasta Adapter file which specifies which adapters to cut out? it does not appear that the -b, -g -a can cut the 28 different sequences I need trimmed.

    Thank you again in advance

    Leave a comment:


  • GenoMax
    replied
    Originally posted by arcolombo698 View Post
    Hello. Thank you in advance.

    do I just use the universal adapter as the input for cutadapt???
    See this guide: http://onetipperday.blogspot.com/201...torprimer.html

    Leave a comment:


  • arcolombo698
    replied
    Trimming Multiple Adapter Sequences

    Hello. Thank you in advance.

    I am in need to trim 3' end sequences off Read1 and all the sequence that follows the 3' end of read 1.

    So looking at cutadapt --help the -a option seems the right fit.

    From the illumina website, for RNA seq there is a list of roughly 27 indexed adapter sequences.

    Is it true that I need to specify all the adapters in the command line for cutadapt to trim them all?

    I just would like some guidance with how to trim the CHIP Seq RNA seq adapter sequences given below:

    OR
    do I just use the universal adapter as the input for cutadapt???




    FROM ILLUMINA WEBSITE ::::
    TruSeq Universal Adapter
    5’ AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCT
    TruSeq Adapter, Index 1 6
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACATCACGATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 2
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACCGATGTATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 3
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACTTAGGCATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 4
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACTGACCAATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 5
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACACAGTGATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 6
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACGCCAATATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 7
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACCAGATCATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 8
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACACTTGAATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 9
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACGATCAGATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 10
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACTAGCTTATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 11
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACGGCTACATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 12
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACCTTGTAATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 13
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACAGTCAACAATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 14
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACAGTTCCGTATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 15
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACATGTCAGAATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 16
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACCCGTCCCGATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 18 7
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACGTCCGCACATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 19
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACGTGAAACGATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 20
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACGTGGCCTTATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 21
    5’ GATCGGAAGAGCACACGTCTGAACTCCAGTCACGTTTCGGAATCTCGTATGCCGTCTTCTGCTTG
    TruSeq Adapter, Index 22

    Leave a comment:


  • ZoeG
    replied
    Originally posted by mmartin View Post
    Hi, I've just released cutadapt 1.4, which comes with some speed improvements, bug fixes and also new features. See the detailed changelog at http://code.google.com/p/cutadapt/ .
    congratulations!
    Thank you, will try ~

    Leave a comment:

Latest Articles

Collapse

  • SEQadmin2
    New Genomics Technologies Take Aim at Long-Standing Limits
    by SEQadmin2


    Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

    We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
    ...
    09-28-2026, 10:25 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by SEQadmin2, 09-29-2026, 09:51 AM
0 responses
42 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-25-2026, 09:06 AM
0 responses
48 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-23-2026, 11:05 AM
0 responses
39 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-18-2026, 11:37 AM
1 response
51 views
0 reactions
Last Post pekgio
by pekgio
 
Working...