I have put together a tutorial website with four core tutorials on it, RNA-Seq, ChIP-Seq, Genome assembly, and SNP calling that may be of use to you.
This website was created to share bioinformatics tutorials and create a dynamic learning environment that does not become dated, PDF contributions welcome and there are four core tutorials available. We would be interested to get some feedback.
Unconfigured Ad
Collapse
X
-
I am not familiar with tophat, but I have been using bowtie.
Before the mapping process of bowtie, it should index the reference genome using the command such as:
bowtie-build hg19.fasta hg19
then six files (hg19.1.ebwt, hg19.2.ebwt, hg19.3.ebwt, hg19.4.ebwt, hg19.rev.1.ebwt, and hg19.rev.2.ebwt) will be generated.
During the process of mapping, bowtie searched the short reads(usually *.fq) again those six indexed files.
I think the index process and mapping process have been automated by tophat.
In your question, it reports "Couldn't find index files C1_R1_1.fq.* ".
I guess the reference genome should be d_melanogaster_fb5_22
did you forget to specify the reference genome in your command?
tophat -p 8 -G genes.gtf -o C1_R1_thout--library-type=fr-firststrand\ genome C1_R1_1.fq C1_R1_2.fq
hope can help you
Originally posted by Ajayi Oyeyemi View PostHi,
I'm very new to RNA sequencing analysis and I've been using Trapnell et al., 2012 protocol as a guide. I've succesfully downloaded the softwares including the fruitfly genome in "my_rna_seq" directory as advised and I try as much as possible to move files extracted in the Downloads to this directory so that I can have everything in the above mentioned directory.
The challenge came when I wanted to align my RNA-seq reads to the reference genome. After the first command, all I got was an error indicating that "Could not find Bowtie index files C1_R1_1.fq.*" I'm faced with this error monster and I just need a clue to find this monster.
I have the following files in my folder:
bowtie-0.12.8
bowtie-0.12.8-linux-x86_64.zip
C1_R1_thout
C1_R1_thout--library-type=fr-firststrand genome
d_melanogaster_fb5_22.1.ebwt
d_melanogaster_fb5_22.2.ebwt
d_melanogaster_fb5_22.3.ebwt
d_melanogaster_fb5_22.4.ebwt
d_melanogaster_fb5_22.ebwt.zip
d_melanogaster_fb5_22.rev.1.ebwt
d_melanogaster_fb5_22.rev.2.ebwt
Drosophila_melanogaster
Drosophila_melanogaster_Ensembl_BDGP5.25.tar.gz
Ensem
genes.gtf
genome.*.
GSE32038_simulated_fastq_files.tar.gz
GSM794483_C1_R1_1.fq.gz
GSM794483_C1_R1_2.fq.gz
GSM794484_C1_R2_1.fq.gz
GSM794484_C1_R2_2.fq.gz
GSM794485_C1_R3_1.fq.gz
GSM794485_C1_R3_2.fq.gz
GSM794486_C2_R1_1.fq.gz
GSM794486_C2_R1_2.fq.gz
GSM794487_C2_R2_1.fq.gz
GSM794487_C2_R2_2.fq.gz
GSM794488_C2_R3_1.fq.gz
GSM794488_C2_R3_2.fq.gz
make_d_melanogaster_fb5_22.sh
README.txt
tophat_out
and I used this command:
tophat -p 8 -G genes.gtf -o C1_R1_thout--library-type=fr-firststrand\ genome C1_R1_1.fq C1_R1_2.fq
Thanks.
Leave a comment:
-
Hi,
I'm very new to RNA sequencing analysis and I've been using Trapnell et al., 2012 protocol as a guide. I've succesfully downloaded the softwares including the fruitfly genome in "my_rna_seq" directory as advised and I try as much as possible to move files extracted in the Downloads to this directory so that I can have everything in the above mentioned directory.
The challenge came when I wanted to align my RNA-seq reads to the reference genome. After the first command, all I got was an error indicating that "Could not find Bowtie index files C1_R1_1.fq.*" I'm faced with this error monster and I just need a clue to find this monster.
I have the following files in my folder:
bowtie-0.12.8
bowtie-0.12.8-linux-x86_64.zip
C1_R1_thout
C1_R1_thout--library-type=fr-firststrand genome
d_melanogaster_fb5_22.1.ebwt
d_melanogaster_fb5_22.2.ebwt
d_melanogaster_fb5_22.3.ebwt
d_melanogaster_fb5_22.4.ebwt
d_melanogaster_fb5_22.ebwt.zip
d_melanogaster_fb5_22.rev.1.ebwt
d_melanogaster_fb5_22.rev.2.ebwt
Drosophila_melanogaster
Drosophila_melanogaster_Ensembl_BDGP5.25.tar.gz
Ensem
genes.gtf
genome.*.
GSE32038_simulated_fastq_files.tar.gz
GSM794483_C1_R1_1.fq.gz
GSM794483_C1_R1_2.fq.gz
GSM794484_C1_R2_1.fq.gz
GSM794484_C1_R2_2.fq.gz
GSM794485_C1_R3_1.fq.gz
GSM794485_C1_R3_2.fq.gz
GSM794486_C2_R1_1.fq.gz
GSM794486_C2_R1_2.fq.gz
GSM794487_C2_R2_1.fq.gz
GSM794487_C2_R2_2.fq.gz
GSM794488_C2_R3_1.fq.gz
GSM794488_C2_R3_2.fq.gz
make_d_melanogaster_fb5_22.sh
README.txt
tophat_out
and I used this command:
tophat -p 8 -G genes.gtf -o C1_R1_thout--library-type=fr-firststrand\ genome C1_R1_1.fq C1_R1_2.fq
Thanks.
Leave a comment:
-
Hi
When making Exon Junction libraries, you suggest to use Make Splice Junction Fasta (USeq software package). I've read on their web page that it's depreciated and Make Transcriptome (USeq's package as well) should be used instead.
Make Transcriptome gives 2 outputs; transcripts and splices. I guess I shall use only splices file in further step(s). Am I right?
Leave a comment:
-
the seven fasta files by Li et al., 2008
Can anyone help me in giving a link on the seven fasta files by Li et al 2008 used in the RNA-seq tutorial posted just of recent?
Leave a comment:
-
Hi,
I am trying to install miRanalyzer stand alone version. I have built the database and managed to download all required tools and packages. When I tried to run it with the test data provided, I get this message: "Failed to load Main-Class manifest attribute from miRanalyzer.jar
". Just wondering what I am doing wrong
Leave a comment:
-
Thank you for the very nice tutorial, and also for guiding me to the 'how-to-wiki'!!
Leave a comment:
-
Hey everyone,
I am new to NGS and I was trying the DE analysis of Di Arabidopsis pathogen data as mentioned in the tutorial. I followed every step and in the end I got the count of significant genes to be 169. I wanted to know how do I get the file that describes all the significant genes?. Any help will be much appreciated since I am still in the learning phase.
Thanks in advance,
Himanshu.
Leave a comment:
-
Thanks for the nice tutorial
I was looking for such a tutorial for some time. I got it now. Thanks a lot!
Leave a comment:
-
Hi,
I was wondering if there is a typo in the script here:
Did you mean to check the sam file? (untreated1.sam)Code:awk '$3!="*"' untreated1.fa|wc -l wc -l untreated1.fa
How can you see in the fastA file how many reads were aligned?
Thanks
Leave a comment:
-
Hi,
I find the script great and very helpful.
I don't understand the part with the GO enrichment.
I am trying to do it with drosophila genes, which I converted into entrez annotations.
But than I get this:Code:> head(pwf) DEgenes bias.data pwf 43072 0 2002 0.018493247 40191 0 14212 0.044941512 318077 0 493 0.009297322 32941 0 1918 0.017987442 42674 0 1182 0.012241218 42675 0 566 0.009468054
Is it because I can't work with goseq on drosophila, or did I miss something in the workflow?Code:> GO.pvals <- goseq(pwf, "dm3", "refGenes") Fetching GO annotations... Error in getgo(rownames(pwf), genome, id, fetch.cats = test.cats) : Couldn't grab GO categories automatically. Please manually specify
Thanks
Assa
Leave a comment:
-
Hello,
Sorry, but after a while I am not able to figure out how this regular expression works:
new_read_chr_names=gsub("(.*)[T]*\\..*","chr\\1",rname(reads))
and can convert these chromosome names:"10.1-129993255" , "11.1-121843856" ,"1.1-197195432" , "12.1-121257530", "13.1-120284312"
Into this new format:
"chr10", "chr11", "chr1", "chr12", "chr13"
I would be very grateful if someone could give me a more detailed explanation about it because I am not able to understand this regular expression.
Thanks in advance!
Leave a comment:
-
Small code issue
Hi Matt,
First of all, fantastic tutorial--very thorough. Just wanted to let you know of a small error in your code segments in your latest update--a missing terminal ')' in your prostate data example when building a TOC.
toc=data.frame(rep(NA,length(tx_by_gene))
Are you currently adding examples or interested in working on a technical report for other DEG software packages (DEGseq, myrna, etc)? I would be interested in working on such a tutorial or wiki.
Originally posted by MDY View PostHello,
I've written a guide to the analysis of RNA-seq data, for the purpose of differential expression analysis. It currently lives on our internal wiki that can't be viewed outside of our division, although printouts have been used at workshops. It is by no means perfect and very much a work in progress, but a number of people have found it helpful, so I thought it would useful to have it somewhere more publicly accessible.
I've attached a pdf version of the guide, although really what I was hoping was that someone here could suggest somewhere where it could be publicly hosted as a wiki. This area is so multifaceted and fast-moving that the only way such a guide can remain useful is if it can be constantly extended and updated.
If anyone has any suggestions about potential hosting, they can contact me at [email protected]
Cheers
Matt
Update: I've put a few extra things on our local Wiki and seeing as people here seem to be finding this useful I thought I'd post an updated version. I'm also an author on a review paper on Differential Expression using RNA-seq which people who find the guide useful, might also find relevant...
RNA-seq Review
Leave a comment:
Latest Articles
Collapse
-
by SEQadmin2
Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.
We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing...-
Channel: Articles
-
ad_right_rmr
Collapse
News
Collapse
| Topics | Statistics | Last Post | ||
|---|---|---|---|---|
|
Started by SEQadmin2, 09-29-2026, 09:51 AM
|
0 responses
37 views
0 reactions
|
Last Post
by SEQadmin2
09-29-2026, 09:51 AM
|
||
|
Started by SEQadmin2, 09-25-2026, 09:06 AM
|
0 responses
44 views
0 reactions
|
Last Post
by SEQadmin2
09-25-2026, 09:06 AM
|
||
|
Started by SEQadmin2, 09-23-2026, 11:05 AM
|
0 responses
37 views
0 reactions
|
Last Post
by SEQadmin2
09-23-2026, 11:05 AM
|
||
|
Started by SEQadmin2, 09-18-2026, 11:37 AM
|
1 response
51 views
0 reactions
|
Last Post
by pekgio
09-21-2026, 02:04 AM
|
Leave a comment: