Originally posted by sdriscoll
View Post
Seqanswers Leaderboard Ad
Collapse
Announcement
Collapse
No announcement yet.
X
-
-
Originally posted by sdriscoll View Postjust to add to this discussion, i found when using sequencing data from mice it worked best for all of my source references to come from UCSC. i used FASTA files for each chromosome downloaded from UCSC's downloads area to build my Bowtie index and I also used UCSC's table browser to produce the GTF file (which i converted to GFF3 using scripts from seq ontology). only when I had built everything from those sources did i have reliable output files that work straight away with the UCSC browser. in fact, when I used the NCBI reference (and swapped the chromosome names out with UCSC's names) the output from Tophat didn't even align with the genome.
Leave a comment:
-
just to add to this discussion, i found when using sequencing data from mice it worked best for all of my source references to come from UCSC. i used FASTA files for each chromosome downloaded from UCSC's downloads area to build my Bowtie index and I also used UCSC's table browser to produce the GTF file (which i converted to GFF3 using scripts from seq ontology). only when I had built everything from those sources did i have reliable output files that work straight away with the UCSC browser. in fact, when I used the NCBI reference (and swapped the chromosome names out with UCSC's names) the output from Tophat didn't even align with the genome.
Leave a comment:
-
Originally posted by melody View Postas the output above
:A few lines of coverage.wig file are:
track type=bedGraph name="TopHat - read coverage"
gi|29823169|ref|NT_025004.13|Hs18_25160 0 9580 0
gi|29823169|ref|NT_025004.13|Hs18_25160 9580 9655 1
then 9580 has 1 or 0 hit??
No, that is a data column because the output is in bedGraph format.
When you copy and paste with correct chromosome name, it will draw a bedGraph based on the value of the data column.
In this example, it will draw 0 for chr18:0-9580 then draw 1 for chr18:9580-9655.
-Statsteam
Leave a comment:
-
as the output above
:A few lines of coverage.wig file are:
track type=bedGraph name="TopHat - read coverage"
gi|29823169|ref|NT_025004.13|Hs18_25160 0 9580 0
gi|29823169|ref|NT_025004.13|Hs18_25160 9580 9655 1
then 9580 has 1 or 0 hit??
Leave a comment:
-
Thank you simon.
I just started bowtie-build with fasta files containing only chromosome names.
Statsteam
Leave a comment:
-
The gi lines you see are the fasta file headers from the NCBI human assembly. Each chromosome in that assembly comes in a separate file and it is the accession codes for those separate files that you are seeing.
The full header for the first accession you found is:
>gi|29823169|ref|NT_025004.13|Hs18_25160 Homo sapiens chromosome 18 genomic contig, reference assembly
You therefore need to find all the accessions for the different chromosomes and replace them with the corresponding chromosome name.
Alternatively you could edit the original fasta files and change the first lines to just contain a chromosome name, eg:
>chr18
..and then reindex the genome and run tophat again. This should put usable chromosome names into your output files.
Leave a comment:
-
Using TopHat output files with UCSC genome browser
Hi all,
Recently, I ran TopHat with 76bp reads data and got the results (sam, bed, and wig files).
Actual a few lines of my input (fasta file) are:
>HWUSI-EAS366:4:1:4:624#0/1:
CTCNGGATGGAGTACAGTGGTGTGATCATGGCTCACTGTAGNNNNNANCN CNTGGGCGCAAGCNNNNNNNNNCTAN
>HWUSI-EAS366:4:1:4:243#0/1:
CGGNGCCGTTGCTGGTTCTCACACCTTTTAGGTCTGTTCTCNNNNNCNGN TNCGACTCTCTCTNNNNNANNNCCGN
>HWUSI-EAS366:4:1:4:1373#0/1:
GAAAAAACCACCCAGCGGTGATGGCAGCGCGCGTGGGTCCCNNNGNGNGN GGGGCGGGTCGCGCNNNNGNNNCGAN
>HWUSI-EAS366:4:1:4:1672#0/1:
GGGCAGGAAAAAAAGGGAAGANAAAATACTGGGGAAGAAAANNNANCNCN GTTTGGCAGCTCTTNNNNGNNNCAGN
And a few lines of junctions.bed file are:
track name=junctions description="TopHat junctions"
gi|29823169|ref|NT_025004.13|Hs18_25160 9690 19656 JUNC00000001 1 + 9690 19656 255,0,0 2 37,38 0,9928
gi|29823169|ref|NT_025004.13|Hs18_25160 14260 19654 JUNC00000002 2 + 14260 19654 255,0,0 2 57,36 0,5358
gi|29823169|ref|NT_025004.13|Hs18_25160 19701 160104 JUNC00000003 3 + 19701 160104 255,0,0 2 32,66 0,140337
A few lines of coverage.wig file are:
track type=bedGraph name="TopHat - read coverage"
gi|29823169|ref|NT_025004.13|Hs18_25160 0 9580 0
gi|29823169|ref|NT_025004.13|Hs18_25160 9580 9655 1
gi|29823169|ref|NT_025004.13|Hs18_25160 9655 9690 0
Here is the problem.
When I copied and pasted the results (either bed file or wig file), I always got an error and when I change the gi|29823169|ref|NT... part to something like chromosome name, it works.
As you can see from my input file, I don't have gi|29823169|ref|NT... part. I am not sure where the TopHat find such label or reference.
Can someone tell me what gi|29823169|ref|NT... part means and how I can convert these files into the one that UCSC genome brower understands. I think I need to get the actual chromosome names.
Thank you,
Statsteam
Latest Articles
Collapse
-
by seqadmin
In recent years, precision medicine has become a major focus for researchers and healthcare professionals. This approach offers personalized treatment and wellness plans by utilizing insights from each person's unique biology and lifestyle to deliver more effective care. Its advancement relies on innovative technologies that enable a deeper understanding of individual variability. In a joint documentary with our colleagues at Biocompare, we examined the foundational principles of precision...-
Channel: Articles
01-27-2025, 07:46 AM -
ad_right_rmr
Collapse
News
Collapse
Topics | Statistics | Last Post | ||
---|---|---|---|---|
Genetic Mapping of Plasmodium knowlesi Identifies Essential Genes and Drug Resistance Mechanisms
by seqadmin
Started by seqadmin, Yesterday, 09:30 AM
|
0 responses
16 views
0 likes
|
Last Post
by seqadmin
Yesterday, 09:30 AM
|
||
Started by seqadmin, 02-05-2025, 10:34 AM
|
0 responses
28 views
0 likes
|
Last Post
by seqadmin
02-05-2025, 10:34 AM
|
||
Started by seqadmin, 02-03-2025, 09:07 AM
|
0 responses
27 views
0 likes
|
Last Post
by seqadmin
02-03-2025, 09:07 AM
|
||
Started by seqadmin, 01-31-2025, 08:31 AM
|
0 responses
35 views
0 likes
|
Last Post
by seqadmin
01-31-2025, 08:31 AM
|
Leave a comment: