Contrail - a hadoop-based de novo sequence assembler

samanta

Senior Member

Join Date: Feb 2010

Posts: 108
- Share
- Tweet
#1

Contrail - a hadoop-based de novo sequence assembler

09-08-2011, 11:16 AM

Hello all,

Most bioinformatics researchers get stuck with the question of how to buy a computer with enough RAM to process their NGS data, because RAM is very expensive. It is not easy to get approval for buying a $100K computer from the managers, when they think everything can be done using $1K laptop and Microsoft excel.

Internet companies like Google developed algorithms to process terabytes and petabytes of data very rapidly and give users the search results. They use clusters of commodity computers with inexpensive disks (hard drive is cheap) using an approach called MapReduce. MapReduce framework is available for free under Hadoop framework distributed by Apache foundation.

Few months back, I came across a genome assembly program called 'contrail' that uses Hadoop to assemble large quantities of NGS data, and it is scalable. When I speak to bioinformaticians about trying out Hadoop instead of buying large and expensive RAM-based machine, I usually hit a hard wall, because the words like Hadoop, MapReduce etc. are foreign to them. So, today I wrote a post to explain setting up and run contrail on your own machine using Hadoop. It is written in such a way that even if you never used Hadoop etc., you can mechanically execute the steps and will be able to assemble the reads in test library in a short time in your own Windows or Unix box. I am hoping that once researchers start to feel that Hadoop approach is easy and scalable for large data sets, they will be able to develop their own programs and the whole community will benefit.

This post discusses how to use contrail assembler -

404 Not Found

http://www.homolog.us/blogs/2011/09/08/contrail-a-de-bruijn-genome-assembler-that-uses-hadoop/

This post discusses how to set up and run Hadoop for a simple sequence analysis example -

404 Not Found

http://www.homolog.us/blogs/2011/08/31/using-hadoop-for-transcriptomics-an-example-to-get-started/

Please note that I am not associated with the researchers, who wrote contrail, and never spoke to them or met them. It is the only example I found for de Bruijn assemblers and decided to try it out.

http://homolog.us
Tags: None

Previous template Next

Exploring the Dynamics of the Tumor Microenvironment

by seqadmin

The complexity of cancer is clearly demonstrated in the diverse ecosystem of the tumor microenvironment (TME). The TME is made up of numerous cell types and its development begins with the changes that happen during oncogenesis. “Genomic mutations, copy number changes, epigenetic alterations, and alternative gene expression occur to varying degrees within the affected tumor cells,” explained Andrea O’Hara, Ph.D., Strategic Technical Specialist at Azenta. “As...
- Channel: Articles
07-08-2024, 03:19 PM

Topics	Statistics	Last Post
Gene Misexpression in the Healthy Human Population by seqadmin Started by seqadmin, Yesterday, 06:46 AM	0 responses 9 views 0 likes	Last Post by seqadmin Yesterday, 06:46 AM
New Method for Rapid Genetic Diagnosis of Mendelian Disorders by seqadmin Started by seqadmin, 07-24-2024, 11:09 AM	0 responses 24 views 0 likes	Last Post by seqadmin 07-24-2024, 11:09 AM
Advancing Nanopore Technology for Portable Sensing Devices by seqadmin Started by seqadmin, 07-19-2024, 07:20 AM	0 responses 159 views 0 likes	Last Post by seqadmin 07-19-2024, 07:20 AM
New RNA-Based Gene Writing Technology Achieves Precise Gene Integration by seqadmin Started by seqadmin, 07-16-2024, 05:49 AM	0 responses 127 views 0 likes	Last Post by seqadmin 07-16-2024, 05:49 AM

Seqanswers Leaderboard Ad

Announcement

Contrail - a hadoop-based de novo sequence assembler

Latest Articles

ad_right_rmr

News