Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • Illumina-Tag sequencing, filtering homopolymers/entropy based filtering

    Hi all, I have a general question to anyone that has been involved in tag-sequencing of functional genes in environmental samples. I use a data processing pipeline based on Usearch (Rob Edgar's methods) and QIIME to analyze functional genes related to the nitrogen cycle in soils and marine environments . I am using usearch, specifically the fastq_stats option, to assess quality values of my sequences over read-length. After visusal inspection of read-length vs. Q value, I am then trimming the sequences to a fixed position, and using QIIME to filter reads using split_libraries_fastq.py with these quality options (-q 19 -r 3 -p 0.75 -n 0). I am using QIIME rather than usearch exclusively because I like to then split the cleaned file with split_fasta_on_sample_ids.py (i have lots of samples, some from different environments, that I would like to analyze independently). Anyway, I have found that for some target genes of lower abundance organisms, there are often a substantial proportion of sequences still in the cleaned files with homo-polymer runs of AAAAAAA, or something similar to this. After doing some investigation, I found that there are ways to eliminate these sequenes using DUST filters, such as those found in the PrinSeq package. However, I am wondering if anyone knows of methods, scripts, or filter options within QIIME or usearch that will remove these sequences. The reason I am concerned about these files is obvious, I hope, and a substantial amount of these reads ends up contributing to my OTU files, yet display no hits to reference gene databases of my targets or even GenBank!

    So, overall, does anyone know if entropy based filters exist in QIIME or usearch for elimination of low-complexity sequences.

    Hope my question makes sense!

    Thanks,

    -Tony

Latest Articles

Collapse

  • seqadmin
    Best Practices for Single-Cell Sequencing Analysis
    by seqadmin



    While isolating and preparing single cells for sequencing was historically the bottleneck, recent technological advancements have shifted the challenge to data analysis. This highlights the rapidly evolving nature of single-cell sequencing. The inherent complexity of single-cell analysis has intensified with the surge in data volume and the incorporation of diverse and more complex datasets. This article explores the challenges in analysis, examines common pitfalls, offers...
    06-06-2024, 07:15 AM
  • seqadmin
    Latest Developments in Precision Medicine
    by seqadmin



    Technological advances have led to drastic improvements in the field of precision medicine, enabling more personalized approaches to treatment. This article explores four leading groups that are overcoming many of the challenges of genomic profiling and precision medicine through their innovative platforms and technologies.

    Somatic Genomics
    “We have such a tremendous amount of genetic diversity that exists within each of us, and not just between us as individuals,”...
    05-24-2024, 01:16 PM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by seqadmin, Today, 07:24 AM
0 responses
4 views
0 likes
Last Post seqadmin  
Started by seqadmin, Yesterday, 08:58 AM
0 responses
11 views
0 likes
Last Post seqadmin  
Started by seqadmin, 06-12-2024, 02:20 PM
0 responses
15 views
0 likes
Last Post seqadmin  
Started by seqadmin, 06-07-2024, 06:58 AM
0 responses
182 views
0 likes
Last Post seqadmin  
Working...
X