Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • adamrork
    replied
    Originally posted by GenoMax View Post
    Rather than doing the filtering you should consider normalizing your data. For this BBMap has another program called bbnorm.sh. You can find a guide here.
    Ah, great point! I'll do that. Thank you!

    Leave a comment:


  • GenoMax
    replied
    Rather than doing the filtering you should consider normalizing your data. For this BBMap has another program called bbnorm.sh. You can find a guide here.

    Leave a comment:


  • adamrork
    started a topic Trimming stringency w/ excessive data

    Trimming stringency w/ excessive data

    I'm in the process of assembling a 500Mb animal genome and have both Illumina and PacBio data to work with using a hybrid assembly approach. Generally when I assemble de novo transcriptomes or genomes, I try to not be too "aggressive" with my quality trimming parameters for raw Illumina reads, running something along the lines of:

    Code:
    bbduk.sh in=file1.fq in2=file2.fq out=trimmed.fq qtrim=rl trimq=10 (plus some adapter trimming parameters)
    However, we sequenced our Illumina libraries incredibly deep and have something on the order of 320x coverage untrimmed and ~270x coverage trimmed using the parameters above (plus some adapter trimming). This is still about twice as much coverage as I would ever generally use for genome assembly. Right now I'm subsetting my trimmed reads to about 100x coverage for hybrid assembly with my PacBio data. I was wondering what folks think about making my quality trimming parameters more stringent given that I have so much excess data to have a higher-quality set of reads on average. For example,

    Code:
    bbduk.sh in=file1.fq in2=file2.fq out=trimmed.fq qtrim=rl trimq=15
    This brings me down to a little over 200x coverage, which I would still subset to 100x coverage-worth of reads from using reformat.sh. I know that this isn't generally recommended, but I've never been quite sure if this is because of lost coverage (which wouldn't be an issue here) or because of something inherently unusual about how higher-quality reads are handled by assemblers. Thus far, my assembly results using the "qtrim=rl trimq=10" parameters seem reasonable, I'm mostly just curious.
    Last edited by adamrork; 12-12-2019, 08:49 PM.

Latest Articles

Collapse

  • seqadmin
    Essential Discoveries and Tools in Epitranscriptomics
    by seqadmin




    The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist...
    04-22-2024, 07:01 AM
  • seqadmin
    Current Approaches to Protein Sequencing
    by seqadmin


    Proteins are often described as the workhorses of the cell, and identifying their sequences is key to understanding their role in biological processes and disease. Currently, the most common technique used to determine protein sequences is mass spectrometry. While still a valuable tool, mass spectrometry faces several limitations and requires a highly experienced scientist familiar with the equipment to operate it. Additionally, other proteomic methods, like affinity assays, are constrained...
    04-04-2024, 04:25 PM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by seqadmin, Today, 11:49 AM
0 responses
11 views
0 likes
Last Post seqadmin  
Started by seqadmin, Yesterday, 08:47 AM
0 responses
16 views
0 likes
Last Post seqadmin  
Started by seqadmin, 04-11-2024, 12:08 PM
0 responses
61 views
0 likes
Last Post seqadmin  
Started by seqadmin, 04-10-2024, 10:19 PM
0 responses
60 views
0 likes
Last Post seqadmin  
Working...
X