Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • shocker8786
    Member
    • Jan 2013
    • 28

    #1

    Error when using HTSeq count

    I used Tophat to produce a bam file from paired end RNAseq data, as well as unpaired reads that were retained after trimming. I then sorted by read name and converted the bam to a sam file. When I run HTSeq with the following arguments:

    python count.py -m intersection-strict -s yes 13_accepted_hits.sam Sus_scrofa.gtf

    I get an error message:

    100000 GFF lines processed.
    200000 GFF lines processed.
    300000 GFF lines processed.
    400000 GFF lines processed.
    496967 GFF lines processed.
    Error occured when processing SAM input (line 4684 of file 13_acc
    epted_hits.sam):
    'pair_alignments' needs a sequence of paired-end alignments
    [Exception type: ValueError, raised in __init__.py:612]

    It looks to me that there are unpaired reads in the sam file that are causing the problem. Do I have to run separate Tophat alignments with the paired-end and unpaired data, then run the sam files through HTSeq and add the resulting counts together? Is there no way to use HTSeq with a both paired and unpaired reads in the same file? Thanks.
  • Jeremy
    Senior Member
    • Nov 2009
    • 190

    #2
    HTSeq requires that the file is name sorted, but the default sorting of some programs is position sorting. I had the same problem, resorting by name fixed it. I think it can handle paired end and non paired end in the same file (not certain), the problem comes when the name indicates a paired read but the next line is a diferent read.

    samtools sort -n in.bam out.sort

    Comment

    • shocker8786
      Member
      • Jan 2013
      • 28

      #3
      As I said in the original post, I sorted by read name. However I know the unpaired reads have the /1 and /2 read name endings, so maybe that could be the issue?

      Comment

      Latest Articles

      Collapse

      • SEQadmin2
        Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
        by SEQadmin2



        CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

        Despite this, “CRISPR helped turn genome editing from a specialized technique into
        ...
        07-31-2026, 11:01 AM
      • SEQadmin2
        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
        by SEQadmin2


        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

        The systematic characterization of the human proteome has
        ...
        07-20-2026, 11:48 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by SEQadmin2, Yesterday, 10:35 AM
      0 responses
      9 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 08-06-2026, 07:41 AM
      0 responses
      27 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 08-03-2026, 10:13 AM
      0 responses
      45 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-31-2026, 02:55 AM
      0 responses
      48 views
      0 reactions
      Last Post SEQadmin2  
      Working...