Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts

  • kenietz
    replied
    @ulz_peter:
    Thank you very much for the effort of compiling that manual and for updating it as well! It is really useful. While going through it I found out i was on the right track

    Leave a comment:


  • Orr Shomroni
    replied
    Would it also be a good idea to trim of the Illumina barcode sequences? And to keep only the barcoded reads?
    I agree with ulz_peter, I don't think the barcodes are supposed to appear in the reads. Regarding Sickle, you cannot determine specificaly which region you want to trim. The only things you can control are the quality threshold (the read is trimmed before and after bases which have a quality below the threshold) and length threshold (if the trimmed read is shorter than the threshold, it will be removed)

    Leave a comment:


  • ulz_peter
    replied
    in case the barcodes are still in your sequence you need to trim them out, they could cause serious trouble in downstream analyses. However, for our MiSeq we get the index reads as separate fastq files, are you sure they are in your sequence?

    Leave a comment:


  • ddaneels
    replied
    Trim barcodes

    Would it also be a good idea to trim of the Illumina barcode sequences? And to keep only the barcoded reads?

    Leave a comment:


  • Orr Shomroni
    replied
    ddaneels, for clipping the 3' end I use Sickle (https://github.com/najoshi/sickle). It is generally used for read trimming, as reads tend to have a low quality on the 3' end and the 5' end, so this tool allows trimming based on the quality from the fastq file.
    Regarding the reference file, I usually use a single FASTA file as the reference which includes all chromosomes, instead of per chromosome. It takes more time to map for the entire genome, but unless you want to study a specific region of your read data, it makes more sense to use the complete reference as an exploratory analysis.

    Leave a comment:


  • ulz_peter
    replied
    That's a good question. Those files correspnd to contigs that have not been reliably placed on a chromosome due to problematic sequences. I personally just use chr1-22 X, Y and M. There should be some chrUn_... contigs as well, they have not been associated with a chromosome so far. It depends a bit what you want to do...
    However, I am not sure if GATK is very happy with those chromosome files. I never tried it out.

    Another thing to consider would be the alternative haplotypes for thr MHC and the other two regions. I am not sure how much they deviate from the original reference sequence, or if there is a good reason for including them.

    Lots of things to consider, but for getting started I'd recommend to use only the chr1-22, X, Y and M(T) chromosomes (in this order)

    Leave a comment:


  • ddaneels
    replied
    Thanks for the fast reply!

    I do have another question: I downloaded the newest UCSC release of the human genome. 7When I unpack this .tar.gz file there are a lot of different files called f.e. chr17_gl000 ...

    Should I also include these in the single genome reference file?

    Leave a comment:


  • ulz_peter
    replied
    thanks a lot. Up to now, there was no need to remove adaptor sequences from the exome-sequencing data we got. However, for other experiments I was using the cutadapt tool (http://code.google.com/p/cutadapt/). Many people use the FastX-Toolkit, which somehow didn't work for my data (from the MiSeq)...

    Leave a comment:


  • ddaneels
    replied
    Hello,

    Really helpful manual! Nice work!

    I saw that you are working with Illumina data. I was wondering if you clipped of the 3' adaptor sequence? If so, with which tool?

    Leave a comment:


  • cllorens
    replied
    Sehrrot
    Try to see if your bed file has header and if so remove it and try to run the command again. I had the same bug report and doing so it worked succesfully.
    Carlos

    Leave a comment:


  • Orr Shomroni
    replied
    What software do you use to integrate the BED file into the analysis? I'm using samtools, and it works fine.

    Did you get a BED file for the Nimblegen kit from the UCSC website? Can you give me a link to it? I don't know how the BED file from UCSC looks like, but the Nimblegen one starts with this line:

    track name=target_region description="Target Regions"

    Also, it has 2 sections: section 1 indicates positions of enrichment regions, and section 2 indicates complete probe regions (so enrichments regions + a few bases before and after them). So it could be that your software cannot deal with the header in the second section.

    Leave a comment:


  • sehrrot
    replied
    Thanks Orr

    I've tried that with V2 bed file but got an error message

    File associated with name ../../apps/annotations/Design_Annotation_files/Target_Regions/SeqCap_EZ_Exome_v2.bed is malformed: Couldn't parse line 'track name=target_region description="Target Regions
    I cannot exactly understand why the format of bed files is different - ucsc bed file and nimblegen one.

    Leave a comment:


  • Orr Shomroni
    replied
    sehrrot, you can also download it from the original Nimblegen website:

    Just open "Design and annotation files", and version 2 of the BED file should be there for download. Make sure also that the Nimblegen kit is the same, because they recently started selling version 3

    Leave a comment:


  • sehrrot
    replied
    thanks for the great manual.

    I have a short question:
    I'm analysing Nimblegen exome and not sure which bed file I have to use for SNP calling. The bed file from nimblegen is not proper format compared to the bed file that Broad inst. provides. Otherwise, can I just use the bed file that the manual recommends (10bp interval target bed file from ucsc)?

    Leave a comment:


  • cllorens
    replied
    The GS ftp site of GATK is just exactly what i needed. thank you Jon.

    Leave a comment:

Latest Articles

Collapse

  • SEQadmin2
    Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
    by SEQadmin2



    CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

    Despite this, “CRISPR helped turn genome editing from a specialized technique into
    ...
    07-31-2026, 11:01 AM
  • SEQadmin2
    Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
    by SEQadmin2


    Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

    The systematic characterization of the human proteome has
    ...
    07-20-2026, 11:48 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by SEQadmin2, 08-11-2026, 10:35 AM
0 responses
11 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 08-06-2026, 07:41 AM
0 responses
30 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 08-03-2026, 10:13 AM
0 responses
48 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 07-31-2026, 02:55 AM
0 responses
48 views
0 reactions
Last Post SEQadmin2  
Working...