Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts

  • BobbyKing
    replied
    I have put together a tutorial website with four core tutorials on it, RNA-Seq, ChIP-Seq, Genome assembly, and SNP calling that may be of use to you.

    This website was created to share bioinformatics tutorials and create a dynamic learning environment that does not become dated, PDF contributions welcome and there are four core tutorials available. We would be interested to get some feedback.

    Leave a comment:


  • hanifk
    replied
    I am not familiar with tophat, but I have been using bowtie.

    Before the mapping process of bowtie, it should index the reference genome using the command such as:
    bowtie-build hg19.fasta hg19
    then six files (hg19.1.ebwt, hg19.2.ebwt, hg19.3.ebwt, hg19.4.ebwt, hg19.rev.1.ebwt, and hg19.rev.2.ebwt) will be generated.

    During the process of mapping, bowtie searched the short reads(usually *.fq) again those six indexed files.

    I think the index process and mapping process have been automated by tophat.

    In your question, it reports "Couldn't find index files C1_R1_1.fq.* ".
    I guess the reference genome should be d_melanogaster_fb5_22

    did you forget to specify the reference genome in your command?
    tophat -p 8 -G genes.gtf -o C1_R1_thout--library-type=fr-firststrand\ genome C1_R1_1.fq C1_R1_2.fq

    hope can help you



    Originally posted by Ajayi Oyeyemi View Post
    Hi,
    I'm very new to RNA sequencing analysis and I've been using Trapnell et al., 2012 protocol as a guide. I've succesfully downloaded the softwares including the fruitfly genome in "my_rna_seq" directory as advised and I try as much as possible to move files extracted in the Downloads to this directory so that I can have everything in the above mentioned directory.
    The challenge came when I wanted to align my RNA-seq reads to the reference genome. After the first command, all I got was an error indicating that "Could not find Bowtie index files C1_R1_1.fq.*" I'm faced with this error monster and I just need a clue to find this monster.
    I have the following files in my folder:
    bowtie-0.12.8
    bowtie-0.12.8-linux-x86_64.zip
    C1_R1_thout
    C1_R1_thout--library-type=fr-firststrand genome
    d_melanogaster_fb5_22.1.ebwt
    d_melanogaster_fb5_22.2.ebwt
    d_melanogaster_fb5_22.3.ebwt
    d_melanogaster_fb5_22.4.ebwt
    d_melanogaster_fb5_22.ebwt.zip
    d_melanogaster_fb5_22.rev.1.ebwt
    d_melanogaster_fb5_22.rev.2.ebwt
    Drosophila_melanogaster
    Drosophila_melanogaster_Ensembl_BDGP5.25.tar.gz
    Ensem
    genes.gtf
    genome.*.
    GSE32038_simulated_fastq_files.tar.gz
    GSM794483_C1_R1_1.fq.gz
    GSM794483_C1_R1_2.fq.gz
    GSM794484_C1_R2_1.fq.gz
    GSM794484_C1_R2_2.fq.gz
    GSM794485_C1_R3_1.fq.gz
    GSM794485_C1_R3_2.fq.gz
    GSM794486_C2_R1_1.fq.gz
    GSM794486_C2_R1_2.fq.gz
    GSM794487_C2_R2_1.fq.gz
    GSM794487_C2_R2_2.fq.gz
    GSM794488_C2_R3_1.fq.gz
    GSM794488_C2_R3_2.fq.gz
    make_d_melanogaster_fb5_22.sh
    README.txt
    tophat_out

    and I used this command:
    tophat -p 8 -G genes.gtf -o C1_R1_thout--library-type=fr-firststrand\ genome C1_R1_1.fq C1_R1_2.fq

    Thanks.

    Leave a comment:


  • Ajayi Oyeyemi
    replied
    Hi,
    I'm very new to RNA sequencing analysis and I've been using Trapnell et al., 2012 protocol as a guide. I've succesfully downloaded the softwares including the fruitfly genome in "my_rna_seq" directory as advised and I try as much as possible to move files extracted in the Downloads to this directory so that I can have everything in the above mentioned directory.
    The challenge came when I wanted to align my RNA-seq reads to the reference genome. After the first command, all I got was an error indicating that "Could not find Bowtie index files C1_R1_1.fq.*" I'm faced with this error monster and I just need a clue to find this monster.
    I have the following files in my folder:
    bowtie-0.12.8
    bowtie-0.12.8-linux-x86_64.zip
    C1_R1_thout
    C1_R1_thout--library-type=fr-firststrand genome
    d_melanogaster_fb5_22.1.ebwt
    d_melanogaster_fb5_22.2.ebwt
    d_melanogaster_fb5_22.3.ebwt
    d_melanogaster_fb5_22.4.ebwt
    d_melanogaster_fb5_22.ebwt.zip
    d_melanogaster_fb5_22.rev.1.ebwt
    d_melanogaster_fb5_22.rev.2.ebwt
    Drosophila_melanogaster
    Drosophila_melanogaster_Ensembl_BDGP5.25.tar.gz
    Ensem
    genes.gtf
    genome.*.
    GSE32038_simulated_fastq_files.tar.gz
    GSM794483_C1_R1_1.fq.gz
    GSM794483_C1_R1_2.fq.gz
    GSM794484_C1_R2_1.fq.gz
    GSM794484_C1_R2_2.fq.gz
    GSM794485_C1_R3_1.fq.gz
    GSM794485_C1_R3_2.fq.gz
    GSM794486_C2_R1_1.fq.gz
    GSM794486_C2_R1_2.fq.gz
    GSM794487_C2_R2_1.fq.gz
    GSM794487_C2_R2_2.fq.gz
    GSM794488_C2_R3_1.fq.gz
    GSM794488_C2_R3_2.fq.gz
    make_d_melanogaster_fb5_22.sh
    README.txt
    tophat_out

    and I used this command:
    tophat -p 8 -G genes.gtf -o C1_R1_thout--library-type=fr-firststrand\ genome C1_R1_1.fq C1_R1_2.fq

    Thanks.

    Leave a comment:


  • Gateway
    replied
    Hi

    When making Exon Junction libraries, you suggest to use Make Splice Junction Fasta (USeq software package). I've read on their web page that it's depreciated and Make Transcriptome (USeq's package as well) should be used instead.

    Make Transcriptome gives 2 outputs; transcripts and splices. I guess I shall use only splices file in further step(s). Am I right?

    Leave a comment:


  • Ajayi Oyeyemi
    replied
    the seven fasta files by Li et al., 2008

    Can anyone help me in giving a link on the seven fasta files by Li et al 2008 used in the RNA-seq tutorial posted just of recent?

    Leave a comment:


  • rna_follower
    replied
    Hi,

    I am trying to install miRanalyzer stand alone version. I have built the database and managed to download all required tools and packages. When I tried to run it with the test data provided, I get this message: "Failed to load Main-Class manifest attribute from miRanalyzer.jar
    ". Just wondering what I am doing wrong

    Leave a comment:


  • lilletine
    replied
    Thank you for the very nice tutorial, and also for guiding me to the 'how-to-wiki'!!

    Leave a comment:


  • himanshu04
    replied
    Hey everyone,
    I am new to NGS and I was trying the DE analysis of Di Arabidopsis pathogen data as mentioned in the tutorial. I followed every step and in the end I got the count of significant genes to be 169. I wanted to know how do I get the file that describes all the significant genes?. Any help will be much appreciated since I am still in the learning phase.
    Thanks in advance,
    Himanshu.

    Leave a comment:


  • caoyh
    replied
    Thanks very much.
    This is the right tutorial for me..

    Leave a comment:


  • yqxiong
    replied
    Thanks for the nice tutorial

    I was looking for such a tutorial for some time. I got it now. Thanks a lot!

    Leave a comment:


  • frymor
    replied
    Hi,

    I was wondering if there is a typo in the script here:
    Code:
    awk '$3!="*"' untreated1.fa|wc -l
    wc -l untreated1.fa
    Did you mean to check the sam file? (untreated1.sam)
    How can you see in the fastA file how many reads were aligned?

    Thanks

    Leave a comment:


  • frymor
    replied
    Hi,

    I find the script great and very helpful.
    I don't understand the part with the GO enrichment.

    I am trying to do it with drosophila genes, which I converted into entrez annotations.
    Code:
    > head(pwf)
           DEgenes bias.data         pwf
    43072        0      2002 0.018493247
    40191        0     14212 0.044941512
    318077       0       493 0.009297322
    32941        0      1918 0.017987442
    42674        0      1182 0.012241218
    42675        0       566 0.009468054
    But than I get this:
    Code:
    > GO.pvals <- goseq(pwf, "dm3", "refGenes")
    Fetching GO annotations...
    Error in getgo(rownames(pwf), genome, id, fetch.cats = test.cats) : 
      Couldn't grab GO categories automatically.  Please manually specify
    Is it because I can't work with goseq on drosophila, or did I miss something in the workflow?

    Thanks
    Assa

    Leave a comment:


  • adansonia
    replied
    Hello,

    Sorry, but after a while I am not able to figure out how this regular expression works:

    new_read_chr_names=gsub("(.*)[T]*\\..*","chr\\1",rname(reads))

    and can convert these chromosome names:"10.1-129993255" , "11.1-121843856" ,"1.1-197195432" , "12.1-121257530", "13.1-120284312"

    Into this new format:

    "chr10", "chr11", "chr1", "chr12", "chr13"

    I would be very grateful if someone could give me a more detailed explanation about it because I am not able to understand this regular expression.

    Thanks in advance!

    Leave a comment:


  • marcel777
    replied
    Small code issue

    Hi Matt,
    First of all, fantastic tutorial--very thorough. Just wanted to let you know of a small error in your code segments in your latest update--a missing terminal ')' in your prostate data example when building a TOC.

    toc=data.frame(rep(NA,length(tx_by_gene))

    Are you currently adding examples or interested in working on a technical report for other DEG software packages (DEGseq, myrna, etc)? I would be interested in working on such a tutorial or wiki.

    Originally posted by MDY View Post
    Hello,

    I've written a guide to the analysis of RNA-seq data, for the purpose of differential expression analysis. It currently lives on our internal wiki that can't be viewed outside of our division, although printouts have been used at workshops. It is by no means perfect and very much a work in progress, but a number of people have found it helpful, so I thought it would useful to have it somewhere more publicly accessible.

    I've attached a pdf version of the guide, although really what I was hoping was that someone here could suggest somewhere where it could be publicly hosted as a wiki. This area is so multifaceted and fast-moving that the only way such a guide can remain useful is if it can be constantly extended and updated.

    If anyone has any suggestions about potential hosting, they can contact me at [email protected]

    Cheers

    Matt

    Update: I've put a few extra things on our local Wiki and seeing as people here seem to be finding this useful I thought I'd post an updated version. I'm also an author on a review paper on Differential Expression using RNA-seq which people who find the guide useful, might also find relevant...

    RNA-seq Review

    Leave a comment:


  • jayangshu.saha
    replied
    Excellent post! Many thanks for sharing.

    Leave a comment:

Latest Articles

Collapse

  • SEQadmin2
    New Genomics Technologies Take Aim at Long-Standing Limits
    by SEQadmin2


    Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

    We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
    ...
    09-28-2026, 10:25 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by SEQadmin2, 09-29-2026, 09:51 AM
0 responses
37 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-25-2026, 09:06 AM
0 responses
44 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-23-2026, 11:05 AM
0 responses
37 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-18-2026, 11:37 AM
1 response
51 views
0 reactions
Last Post pekgio
by pekgio
 
Working...