Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • JWolters
    Member
    • Jan 2013
    • 13

    Fasta formatting issues

    I am trying to align a multifasta containing various regions of interest from a reference genome and align them to a contig assembled from pacbio reads generated from a related strain, just to get a picture of what regions are present and the degree of mismatch/rearrangement.

    Ive been trying to use nucmer, but am getting a strange error which I've put into the mummer-help mailing list and am still awaiting a response:
    "postnuc: tigrinc.cc:337: int Read_String(FILE*, char*&, long int&, char*, int): Assertion `Len > 0 && Line [Len - 1] == '\n'' failed."

    The syntax of the error message made me suspect there was a formatting issue with my sequences, but I could find no empty sequences, and no missing newline characters.

    I switched to just trying to run the alignment via blastn with the align 2 sequences option, but received the following error:
    "Message: Message: NCBI C++ Exception:# "local_db_adapter.cpp", line 123: Error: ncbi::blast::s_CheckForBlastSeqSrcErrors() - NCBI C++ Exception:# "blast_setup.hpp", line 190: Error: Sequence contains no data##"

    Once again it appears to be a formatting error in the fasta files, but I cannot find any empty sequences.

    Looking into it, I thought it might be an issue with the line length in the sequences. I've run nucmer before on sequences of 85 kbp all in one line in the fasta file and had no issues before, but I figured I would try it.

    So I transformed the fasta files to only have 80 characters per line in the sequences (max limit I read somewhere per line, but this does not seem be a standard rule).

    These transformed files gave the exact same error messages in both nucmer and nblast.

    I must be missing something, and if these error messages are anything to go by it should be something obvious but I just can't seem to find it.

    Couldn't attach the files due to size, so I've uploaded them here:
    Features, 80 char per line: http://pastebin.com/E1MyBstr
    Features: http://pastebin.com/rfsiiyKJ
    Contig (view as Raw):http://pastebin.com/tyVsEP8g
    Contig, 80 char per line: http://pastebin.com/DmZBCtiv


    Any help would be vastly appreciated!

    P.S. I thought I could use bwa to do this but apparently that is just for aligning fastq reads to a reference?
  • GenoMax
    Senior Member
    • Feb 2008
    • 7142

    #2
    Not answering your question directly ...

    I remember from a past thread that someone had suggested using "mauve" (http://gel.ahabs.wisc.edu/mauve/) for this type of analysis.

    Comment

    • JWolters
      Member
      • Jan 2013
      • 13

      #3
      After looking into this, Mauve really might be just the right tool for the job since it can apparently export sets of positionally orthologous features (genes, CDS, tRNA, and so on), thus getting me all the features in which I was interested.

      Thanks for the advice!

      Comment

      Latest Articles

      Collapse

      • SEQadmin2
        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
        by SEQadmin2


        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

        The systematic characterization of the human proteome has
        ...
        07-20-2026, 11:48 AM
      • SEQadmin2
        Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
        by SEQadmin2



        Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
        ...
        07-09-2026, 11:10 AM
      • SEQadmin2
        Cancer Drug Resistance: The Lingering Barrier to Rising Survival
        by SEQadmin2



        Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

        There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
        07-08-2026, 05:17 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by SEQadmin2, Yesterday, 12:17 PM
      0 responses
      13 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-23-2026, 11:41 AM
      0 responses
      15 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-20-2026, 11:10 AM
      0 responses
      23 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-13-2026, 10:26 AM
      0 responses
      37 views
      0 reactions
      Last Post SEQadmin2  
      Working...