Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • smylalwys
    Junior Member
    • Jan 2015
    • 4

    #1

    Hadoop for human genome data

    Hello Everyone,

    How do we store the human genome data using Hadoop (chromosome level) so that we can perform processing (bio-algorithm computing) on the data using Hadoop clusters?
  • dpryan
    Devon Ryan
    • Jul 2011
    • 3478

    #2
    How one best stores the data is entirely dependent on how the actual cluster is constructed and what the nature of the algorithm is. If the cluster is essentially a cloud with slow IO then you'll approach this differently than with a HPC cluster with a faster local storage array. Also, if you just need to load the genome into memory for long computations then it doesn't really matter how you store it, that's not going to be the bottleneck.

    Comment

    • smylalwys
      Junior Member
      • Jan 2015
      • 4

      #3
      Hi Ryan,
      Thanks for your reply. We do have cluster of 30 machines with hadoop. The problem is we are planning to process the human genome project using hadoop. Here the data is in the form of BAM files. I know if I load the data to hdfs, it will automatically split it into chunks and store on the name nodes. Thats is the problem here. I couldn't split the data like that. Need to split the data chromosome wise so that we can perform bio algorithm computing on them.

      Can someone please give some insights on this

      Comment

      • dpryan
        Devon Ryan
        • Jul 2011
        • 3478

        #4
        Without knowing more detail it's impossible to give any guidance. Hadoop is a general tool to facilitate processing. How you should split things depends entirely on what you want to do with the results (and "bio algorithm computing" has absolutely no meaning).

        Comment

        • smylalwys
          Junior Member
          • Jan 2015
          • 4

          #5
          Bio algorithm computing : for instance bisulfite methylation extraction

          Comment

          • dpryan
            Devon Ryan
            • Jul 2011
            • 3478

            #6
            Yes, that's one of many possible but completely unrelated tasks. I've already responded to this on one of your biostars threads. Please don't cross post.

            Comment

            • smylalwys
              Junior Member
              • Jan 2015
              • 4

              #7
              Currently we use bismap ( python tool ). Is there a way to store the data chromosome wise on hadoop.and run the bismap tool command as map reduce jobs

              Comment

              • dpryan
                Devon Ryan
                • Jul 2011
                • 3478

                #8
                How this would be done would depend entirely on the cluster, but there's generally no single command (or simple series thereof) that would allow that. The traditional way to do this would be to simply tell BSMap's methylation extractor to just process a single chromosome (and then run that simultaneously with different chromosomes on different cores). You could simply do that in a fraction of the time it's take to implement a full hadoop-based solution.

                Comment

                Latest Articles

                Collapse

                • SEQadmin2
                  Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                  by SEQadmin2



                  CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                  Despite this, “CRISPR helped turn genome editing from a specialized technique into
                  ...
                  07-31-2026, 11:01 AM
                • SEQadmin2
                  Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                  by SEQadmin2


                  Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                  The systematic characterization of the human proteome has
                  ...
                  07-20-2026, 11:48 AM

                ad_right_rmr

                Collapse

                News

                Collapse

                Topics Statistics Last Post
                Started by SEQadmin2, 08-13-2026, 12:22 PM
                0 responses
                29 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 08-11-2026, 10:35 AM
                0 responses
                24 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 08-06-2026, 07:41 AM
                0 responses
                38 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 08-03-2026, 10:13 AM
                0 responses
                51 views
                0 reactions
                Last Post SEQadmin2  
                Working...