Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • prasadg
    Member
    • Mar 2012
    • 16

    Dwgsim reads not generating header properly

    Hello,

    I have been using dwgsim to stimulate the reads for my masters thesis project.
    reads generated by dwgsim are as follows:

    @gi|29366675|ref|NC_000866.4|_148598_1_0_1_0_0_7:0:0_0:0:0_0/1
    TCAATGTTTAAAATTTTTCTATAAGCCTTTAACTTTATAGAATTATTATTCCAGACTAAATTTTCAGTCTGTTCATCATGTTTATCAATGATATTTAAAATCGAATCAAGCAAGATAAACGTCTCAAACGAAATTATGTTCGATTGCAGAAGTTTAAAAATATAACTTGATTGAACTTTTGAATTATACTCAAAAATTTCTTTAAAAGCAGAAACTTCAACTTTTTTACTAAAATAATAAATGTTGCGAATATCTTCTTCAAACTTAAATTTAATTTGCTCTAAGCGTCCGATATATTCA
    +
    274243352322242112304101.421322033226312022224227224543221700163322322254222233202222323410323211235002421417.1212/32032144113/34243333142240265131237222310434142423502121331306222223125411125322651421222221301325123622224320232443510122.240104223//211227532272513317.654312223233312046223331322/3520
    @gi|29366675|ref|NC_000866.4|_76814_1_1_0_0_0_6:0:0_0:0:0_1/1
    TGGACAAGAATTTTTTATTGAAATAAAACCTAAAAAAGAAACACAACCACCAGTTAAACCAGCACATCTAACGACCGCAGCGAAGAAAAGATTTATGAATGAAATTTATACCTGGTCTGTGAACACTGACAAATGGAAAGCAGCACAATCTTTAGCTGAAAAGCGTGGAATAAAATTTAGAAGTCTAACAGAAGATGGATTAGGAGCTCTTGGCTTTAAGGGGGCATAATGGCTATTTTTTAAATAATTAATGAAAGCACTCCCCAATTTCCAAAGGTTAAGCAATCATTAAATGATAAG
    +
    22522022353212222523315421112234/62214021243134223,43/1223252232230232/3.223202222224233232313252311325252022243342030246412026523234522/22132/22/15341151224513230152142411053242222234-22431351303212121442/22222144322524220/251221301223332102013114322254132.252140334236327221301340224342321213223224
    @gi|29366675|ref|NC_000866.4|_107901_1_0_1_0_0_5:0:0_0:0:0_2/1
    TGCTCCTAAATTCTTGGAAGATGTGCTTGCAACTGAAATGGCAGATGAAATCAATAAAGACATTCTGCAGTCTATGATTACAGTGTCAAAACGCTATAAAGTTACAGGAATTACTGATAGTGGATTCATCGATTTGAGTTATGCATCTGCTCCTGAAGCTGGTCGTTCATTATACCGAATGGTATGTGAAATGGTTTCGAATATCCAAAAAGAATCAACTTATACAGCAACGTTCTGTGGTGCTTCAGCTCGTGCCGCTGCGATTCTTGCTGCATCAGGCTGGTTAAAACATAAACCAGA
    +
    000122222226252023235322.522002744/4354861323223512353233012456213104435423/2061742/2303222333236.531/6/22113.2322841222215253302242/04024543251222220232343020061613.32023222362232225221145475/032212202230/44451252162-0423101/3104/64221342214.15423122522032240254130331452422223422/022232323313226243
    @gi|29366675|ref|NC_000866.4|_98258_1_1_0_0_0_2:1:0_0:0:0_3/1
    TGTCCGGAGATAATAAAGTCATTTTTAATCCTCTTTAATATGCTTTAAAATATTTATACCATTGACATACCATGAGATACTGGAACATACTCAGCAGAATGAACCGAATCCACAAATATAACTGGCGCGTAGTCGTCGCTCATATCCTGAAGCTCTTTTGAAAATACTTCAGATGCTAATCGCATGTCATCTTTATCCGCATAATCAATAAATTTTGACTGCGTTGATAACCATCCAAAAATCACTAAAGACATTCCTAAATCGTCATGATAACCTTCTTCAGCCGCCCAAGACACGCCT
    +
    2211223513131223225222123224233522355451323223322312222200121/122332322306232222122122420532154611333313662/202536242223136710122332362322223533020225222221/3/42424432025020224312234203210.41425200660222222261/222523444326015421214235422/430203321424132223//52-2222.3321431123251222/25312263422323023

    As per the documentation of dwgsim given on sourceforge: in reads name header after 2nd underscore there should be start end 2 (zero-based) but in my case its always 1. I dont know what is that happening. Have anybody faced problem like this? Any help is greatly appreciated!!
  • mastal
    Senior Member
    • Mar 2009
    • 666

    #2
    What type of data are you trying to simulate?

    According to source forge
    "The FASTQs for BWA are split into two files, the first file for one end, the second file for the other. For paired end reads, this means that E1 is in the first file and E2 is in the second file."

    Comment

    • prasadg
      Member
      • Mar 2012
      • 16

      #3
      Originally posted by mastal View Post
      What type of data are you trying to simulate?

      According to source forge
      "The FASTQs for BWA are split into two files, the first file for one end, the second file for the other. For paired end reads, this means that E1 is in the first file and E2 is in the second file."

      Thanks alot for reply Mastel,

      I want to create metagenomic reads. I want to create single end reads. So I was using following command:
      dwgsim -C 10 -1 300 -2 0 test.fasta out

      I think this is happening cause I am keepin -2 as 0. but now when i kept -2 an 300 it did give me start end 2.

      I am new to dwgsim. Can you tell me how can i stimulate single end read and get header with start end 2?

      Comment

      • mastal
        Senior Member
        • Mar 2009
        • 666

        #4
        I have used other simulation software, but not dwgsim.

        My understanding, after reading the source forge page, is that you will only get End2 values if you generate paired end reads or mate pair reads.

        Comment

        • prasadg
          Member
          • Mar 2012
          • 16

          #5
          Originally posted by mastal View Post
          I have used other simulation software, but not dwgsim.

          My understanding, after reading the source forge page, is that you will only get End2 values if you generate paired end reads or mate pair reads.
          Oh ok .

          Firstly i thought the end 2 values are where the sequence had ended. So when I am keeping length of first read and second read 300. I am getting following output

          @gi|29366675|ref|NC_000866.4|_33409_33681_0_1_0_0_4:0:0_6:0:0_0/2
          TAATATTAAAACCCTGCAGTCGTTGGCAAATGATATTCGCAATAAAAAGCAATCTCTGATCGCAGCAGTAGATAAAGCTAAAAAAGTTCAAGCGGCTATAGAAAAAGCATCTTCTGAGTTTATTGATCATGCTGATGAAATAGCACTGCTTCAAGAAGAACTTGATAAAATTGTTAAGACAAAAACTAATTTAGTAATGGAAAAATATCACCGAGGAATTTTGACTGATGTGCTCAAAGATTCTGGTATTAAAGGTGCTATTATTAAAGAGTACATTCCATTATTTAATAAGCAGATTAA
          +
          4205034434320-222263232022112521140114136152/232253122241252241522252243213222122126043226243422206623122445422422251/026242223222216442203261213213232/22123362225222321460230262462322252132132422302271331122331333644.22311342613242252273223540422321223122220224/232432424421300722-230425/32124026021
          @gi|29366675|ref|NC_000866.4|_119906_120098_0_1_0_0_2:0:0_4:0:0_1/2
          TTAGAAAATCTAGCAGCAAGTTCTTTTTTAACTGCCGGGGAATTATTTAAATCCGGGTCATCCATCCGTTTTTTAAGGTCTTCTTAGGCAGCTTCAACTGATTTAACCGTTGAGTCTTTACTCATATCAGCTGAATCGGCATATTTTTCAAAACGAATCATCGCAGCTCGAGCTTCATTAGCCTTCATTAAAGCATTTTTTCTTTCTTCCGGTGAAAGTTGCTCTAATTTTTCTTCTTCTGCCGCACGCTCTTCGTCGGTAGTCAGCGCTTCTTTATTATCTACACCACGAATCCAGTTA
          +
          6332230203154523124232221227314264042026331352032620321324531342242321315122063221566440250/164228223405221232201353/2222422241225221323221245322262433433110125120503/33215222111251435116204414233/24123124322525233322345224272232103126132511135642252/55333424642/1742222131312222132221247220224222143
          @gi|29366675|ref|NC_000866.4|_35520_35694_0_1_0_0_3:1:0_8:1:0_2/2
          GTCACTGGGATCTGAATGGATTTTATATTTATAAAGGAATGGAATCTCATGGTCTTGAACCCGATTTCCTTAAGACTTATAAAGAAGTGTGGTCTGGTCATTTCCATACTATTTCTGCGGCTGCAAACGTTAGATATATTGGGACACCATGGACACTAACCGCAGGTGACGAGAATGACCCTCGTGGGTTCTGGATGTTTGATACAGAAACAGAACGAACGGAATTTATCCCAAACAATACTACCTGGCATCGTAGAATTCATTATCCATTTAAAGGAAAAACTGACTATAAAGATTTTC
          +
          32433322324425213201033414534/3423212122342522520234343/45213242215522212224712202132043563/0332322603120513211238/11044233222233211234211/34

          So perception was if end 1 is 33409(from 1 sequence) then if i am taking length of 1st is 300. so end value 2 should 33409+300 = 33709 but where it is 33681. I am confused about what is end 1 and end 2 value?

          If you know then please can you explain me what is end 1 and end 2 values?

          Comment

          • mastal
            Senior Member
            • Mar 2009
            • 666

            #6
            What sequencing platform (e.g. Illumina, Ion Torrent, SOliD) are your simulated reads supposed to be from?

            I think if you read a bit about the sequencing technology of whichever platform you are trying to simulate, you will understand what the reads should be like.

            Comment

            • prasadg
              Member
              • Mar 2012
              • 16

              #7
              Reads are stimulated from Illumina. I will look in to it.

              Thanks alot for all your help!!

              Comment

              Latest Articles

              Collapse

              ad_right_rmr

              Collapse

              News

              Collapse

              Topics Statistics Last Post
              Started by SEQadmin2, Yesterday, 10:09 AM
              0 responses
              9 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 06-04-2026, 08:59 AM
              0 responses
              17 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 06-02-2026, 12:03 PM
              0 responses
              26 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 06-02-2026, 11:40 AM
              0 responses
              21 views
              0 reactions
              Last Post SEQadmin2  
              Working...