Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts

  • para_seq
    replied
    Setting: Number of 'N's allowed in the reads

    Hi, Ben,

    I ran Bowtie for reads from Illumina sequencing. It is a very good pieces of software for its high speed and relatively small memory usage. I have two related questions. 1. I wonder if there is a switch in Bowtie to filter out reads that contain more than certain number of 'N' s. 2. How many 'N's in each read (e.g. 35-nt long) do people usually allow for Illumina data? Thank you.

    Leave a comment:


  • forrest
    replied
    Hi Ben,
    Can Bowtie do DNA methylation aligment? I haven't found in your manual. How to define the parameter?

    Thank you!

    Leave a comment:


  • Ben Langmead
    replied
    Originally posted by juan View Post
    How does Bowtie handle RNA-seq data, has anyone tried? Can it map split reads? Any plans to add split-read mapping functionality?
    Hi Juan,

    Check out TopHat (linked to from the sidebar of the Bowtie website). TopHat was written by Cole Trapnell and it implements a layer on top of Bowtie that handles spliced alignments, along with several other aspects of alignment to the transcriptome and calling junctions.

    Hope that helps.

    Thanks,
    Ben

    Leave a comment:


  • ewilbanks
    replied
    Originally posted by Ben Langmead View Post
    Hi Layla,

    h_sapiens indexes the NCBI human reference contigs and h_sapiens_asm indexes the NCBI human reference assembly. Take a look at the scripts/make_h_sapiens.sh and scripts/make_h_sapiens_asm.sh files distributed with Bowtie to see exactly what fasta files were indexed and how.

    People often prefer the assembly because the coordinates output by bowtie are more immediately useful (e.g., they correspond to the hg18 coordinates in the Genome Browser).

    Thanks,
    Ben
    Hi Ben,

    Thanks for this, I was confused about this as well. Might be a useful tidbit to put near the downloads on your website? Thanks for all the support. Bowtie rocks!

    Lizzy

    Leave a comment:


  • ewilbanks
    replied
    Originally posted by Ben Langmead View Post
    Hi Lizzy,

    I'd expect, oh, about 7-8 hours or so. Did it finish?

    Thanks,
    Ben
    Hi Ben,

    It did! Total run time was 9 hrs 22 min. Thanks!

    Lizzy

    Leave a comment:


  • juan
    replied
    Split Read RNA-Seq mapping with bowtie?

    How does Bowtie handle RNA-seq data, has anyone tried? Can it map split reads? Any plans to add split-read mapping functionality?

    Leave a comment:


  • Ben Langmead
    replied
    Hi Layla,

    h_sapiens indexes the NCBI human reference contigs and h_sapiens_asm indexes the NCBI human reference assembly. Take a look at the scripts/make_h_sapiens.sh and scripts/make_h_sapiens_asm.sh files distributed with Bowtie to see exactly what fasta files were indexed and how.

    People often prefer the assembly because the coordinates output by bowtie are more immediately useful (e.g., they correspond to the hg18 coordinates in the Genome Browser).

    Thanks,
    Ben

    Leave a comment:


  • Layla
    replied
    Im a newbie to Bowtie....tired of the counting down the hours using MAQ.

    Currently building an index using Bowtie. What is the difference between
    h_sapiens_asm.ebwt.zip and
    h_sapiens.ebwt.zip

    Thanks

    L

    Leave a comment:


  • Ben Langmead
    replied
    Hi Lizzy,

    I'd expect, oh, about 7-8 hours or so. Did it finish?

    Thanks,
    Ben

    Leave a comment:


  • ewilbanks
    replied
    Indexing human genome?

    Hi!

    I'm working on building an index of human genome locally and I was wondering how long this usually takes? Its been running for about 3 hrs, just wondering what to expect. I'm on a MAC dual core with 4GB ram.

    Thanks!
    Lizzy

    Leave a comment:


  • Ben Langmead
    replied
    Originally posted by davisc View Post
    I would like to know how Bowtie handles N's in the indices? I am wondering if it is possible to cut down the mapping time by building and mapping against a repeatmasked version of the genome?
    When Bowtie indexes the reference, it elides non-A/C/G/T characters. So if you index a reference with stretches of Ns, Bowtie will never report an alignment spanning any of the stretches.

    And yes, mapping against the repeatmasked version of the genome (and omitting -m 1) ought to be noticeably faster.

    Ben

    Leave a comment:


  • davisc
    replied
    Question about RepeatMasked hg18 index

    I'm doing RNA-Seq on human samples. In many instances I am mapping using the -m1 -v2 --best criteria to the preassembled hg18.asm index available on the download site. I would like to know how Bowtie handles N's in the indices? I am wondering if it is possible to cut down the mapping time by building and mapping against a repeatmasked version of the genome?

    Leave a comment:


  • bioinfosm
    replied
    thanks Ben.. the light bulb just flashed on me!

    Leave a comment:


  • Ben Langmead
    replied
    Hi boinfosm,

    Originally posted by bioinfosm View Post
    The total of leftover and mapped is less than what we started with. Are the remaining reads mapping to multiple locations, and thus omitted in both these files?
    That shouldn't be the case. When only --un is used (as opposed to both --un and --max), both the unaligned reads and the reads with a number of alignments exceeding the -m limit will go into the --un file. But you're not using the -m option, so no reads should be suppressed due to multiple alignments.

    How are you counting the number of reads in your input set? Note that grep -c '^@' isn't necessarily correct because quality strings can also start with @.

    Thanks,
    Ben

    Leave a comment:


  • bioinfosm
    replied
    I wanted to discuss a use-case:
    A collection of 172 million reads ranging from 36 to 76 base long was used with bowtie to map to a reference.

    $ ./bowtie --best --un leftover -p 4 -t reference reads mapped
    $ grep -c '^@' leftover
    154828705
    $ wc -l mapped
    16269083 mapped

    The total of leftover and mapped is less than what we started with. Are the remaining reads mapping to multiple locations, and thus omitted in both these files?

    Leave a comment:

Latest Articles

Collapse

  • SEQadmin2
    New Genomics Technologies Take Aim at Long-Standing Limits
    by SEQadmin2


    Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

    We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
    ...
    09-28-2026, 10:25 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by SEQadmin2, 09-29-2026, 09:51 AM
0 responses
41 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-25-2026, 09:06 AM
0 responses
47 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-23-2026, 11:05 AM
0 responses
38 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-18-2026, 11:37 AM
1 response
51 views
0 reactions
Last Post pekgio
by pekgio
 
Working...