Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • sheremey
    Junior Member
    • Oct 2010
    • 7

    "nucleotide coverage" to genome feature coverage

    Hello,

    I am absolutely new to the area and am doing my self-education based on web resources. I try to establish a pipeline for bacterial RNA-seq and sRNA discovery based on Illumina NGS data. I primarily use R as my working language; this area seems to be at an early stage and Bioconductor packages are scattered. I've already written my own script to count genome coverage at a single-nucleotide level (leveraging functionality of ShortRead and Biostrings packages). I've also tried a conventional route from FASTQ to BAM files using Bowtie + samtools. As a result, here I am: start from FASTQ => Bowtie => SAMtools => BAM files; I also have my single-nucleotide coverage matrices which seem to match the information in the SAM/bam files.

    Here is a point where I suddenly stuck:

    Conversion of a "raw" genome coverage to coverage based on genomic features. I tried to read docs to GenomicRanges, IRanges etc packages but cannot see a clear path through them.

    What is your favorite way from BAM (using genome annotation .gff files) to a matrix of a feature-centric coverage ready to be used for differential expression tests? And again, the genome has nothing to do with Homo sapiens.

    Thanks a lot!
  • sheremey
    Junior Member
    • Oct 2010
    • 7

    #2
    P.S. One part of the Galaxy pipeline does exactly what I need, via 2 of their functions:

    GFF2anno
    eQuant

    However, GFF2anno doesn't recognize GFF files for bacterial genomes (downloaded right from GenBank)!!!

    This conversion (from raw sequence coverage to genomic features-based coverage) is supposed to be a simple short step, and I believe dozens of tools should do this; the question is - what are they?

    Comment

    • dariober
      Senior Member
      • May 2010
      • 311

      #3
      What about the HTseq package for python (http://www-huber.embl.de/users/ander...verview.html)?

      The command htseq-count (http://www-huber.embl.de/users/ander...unt.html#count) seems to do what you need (if you have BAM you also have the SAM file). From the doc page:
      htseq-count [options] <sam_file> <gff_file>

      It's not R but it should do the job.

      Dario

      Comment

      Latest Articles

      Collapse

      • SEQadmin2
        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
        by SEQadmin2


        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

        The systematic characterization of the human proteome has
        ...
        07-20-2026, 11:48 AM
      • SEQadmin2
        Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
        by SEQadmin2



        Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
        ...
        07-09-2026, 11:10 AM
      • SEQadmin2
        Cancer Drug Resistance: The Lingering Barrier to Rising Survival
        by SEQadmin2



        Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

        There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
        07-08-2026, 05:17 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by SEQadmin2, 07-24-2026, 12:17 PM
      0 responses
      21 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-23-2026, 11:41 AM
      0 responses
      19 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-20-2026, 11:10 AM
      0 responses
      26 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-13-2026, 10:26 AM
      0 responses
      38 views
      0 reactions
      Last Post SEQadmin2  
      Working...