Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • puggie
    Member
    • Nov 2011
    • 52

    #1

    Mapping on Windows 7 PC (Galaxy etc.)

    Dear all,

    I would like to map som RNA-seq data using BWA. This can be done in Galaxy for instance. However, I have one problem regarding the reference genome. It is mm9 with one additional custom chromosome (as I'm looking for special occurences of fusion events). The chromosomes I have in FASTA format download from UCSC, however how do I make a multifasta file or even a len file for mapping? I have tried to open in text editors but the apps run out of memory.
  • maubp
    Peter (Biopython etc)
    • Jul 2009
    • 1544

    #2
    If you have to use Windows, you can do quite a lot at the command line - especially if you install Cygwin for a Linux like environment. You could also do sequence manipulation with a scripting language like Perl or Python - both BioPerl and Biopython should be fine under Windows.

    However, in the long run I think you will find sequencing data analysis easier under Linux or Mac OS X (which is a type of Unix) than on Windows - simply because this is what most of the cutting edge tools are designed and tested on.

    Alternatively, sticking with Galaxy, you can concatenate FASTA files together to make a new reference (i.e. combine the mm9 FASTA with your custom chromosome FASTA) using their "Concatenate datasets" tool.

    Comment

    • puggie
      Member
      • Nov 2011
      • 52

      #3
      Thanks for your reply. I think I will stick with galaxy for now as I have only reached the learning phase of data analysis. However, in the long run we may setup a dedicated work station.

      Do you perhaps know, if it is possible to strip down the read lengths of paired end reads? The thing is I have 101 bp reads, is it possible to e.g. strip down to 50 bp and map on basis of that?

      Comment

      • maubp
        Peter (Biopython etc)
        • Jul 2009
        • 1544

        #4
        Originally posted by puggie View Post
        Do you perhaps know, if it is possible to strip down the read lengths of paired end reads? The thing is I have 101 bp reads, is it possible to e.g. strip down to 50 bp and map on basis of that?
        Yes, but why do that? Having longer reads should give more specificity for the mapping. Trimming the reads based on their individual qualities makes more sense - I think there are Galaxy videocast/tutorials on this kind of thing.

        Comment

        • puggie
          Member
          • Nov 2011
          • 52

          #5
          Originally posted by maubp View Post
          Yes, but why do that? Having longer reads should give more specificity for the mapping. Trimming the reads based on their individual qualities makes more sense - I think there are Galaxy videocast/tutorials on this kind of thing.

          Yes I agree. The think is that we are investigating special kinds of chimera, where we expect many reads to catch chimeric sequences. Hence a number of reads will be chimeric. In Galaxy I can apply certain filters as to match read pairs in which each mate map to different chromosomes (mouse chromosome + our custom chromosome). However, will we loose data if say Read 1 is chimeric over 101 bp length?

          It has be noted that many transcripts will start from our custom genome, and read length to custom genome may be below 101 bp.

          That is why I figured I could do the mapping with full length reads, and go shorter and compare.

          Comment

          • maubp
            Peter (Biopython etc)
            • Jul 2009
            • 1544

            #6
            I can see why you might try this now. Have you explored the "Trim (leading or trailing characters)" tool in Galaxy? It can handle FASTQ reads.

            Comment

            • puggie
              Member
              • Nov 2011
              • 52

              #7
              Originally posted by maubp View Post
              I can see why you might try this now. Have you explored the "Trim (leading or trailing characters)" tool in Galaxy? It can handle FASTQ reads.
              I will try out that tool. Im concatenating my genome assembly now. Thx.

              EDIT: One more thing... Can you please tell me maubp the speed you get via ftp://main.g2.bx.psu.edu ?? I'm on 100+ Mbit line however I get around 5Mbit upload.
              Last edited by puggie; 03-08-2012, 05:52 AM.

              Comment

              • maubp
                Peter (Biopython etc)
                • Jul 2009
                • 1544

                #8
                I've never used FTP with the public Galaxy - I've only ever uploaded small files by HTTP or copy&paste.

                Given you are in Europe and the Galaxy server is somewhere in America, getting 5Mbit uploads doesn't sound too bad. Perhaps ask on the galaxy-user mailing list if this is typical?

                Comment

                Latest Articles

                Collapse

                • SEQadmin2
                  Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                  by SEQadmin2



                  CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                  Despite this, “CRISPR helped turn genome editing from a specialized technique into
                  ...
                  Today, 11:01 AM
                • SEQadmin2
                  Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                  by SEQadmin2


                  Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                  The systematic characterization of the human proteome has
                  ...
                  07-20-2026, 11:48 AM
                • SEQadmin2
                  Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                  by SEQadmin2



                  Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                  ...
                  07-09-2026, 11:10 AM

                ad_right_rmr

                Collapse

                News

                Collapse

                Topics Statistics Last Post
                Started by SEQadmin2, Today, 02:55 AM
                0 responses
                7 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-24-2026, 12:17 PM
                0 responses
                11 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-23-2026, 11:41 AM
                0 responses
                12 views
                0 reactions
                Last Post SEQadmin2  
                Started by SEQadmin2, 07-20-2026, 11:10 AM
                0 responses
                24 views
                0 reactions
                Last Post SEQadmin2  
                Working...