Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts

  • IsBeth
    replied
    Thanks Mike for your answer =)

    I should have said before that I have read already the brief explanations about the different methods (sorry), yet my problem is that I have not enough statistic skills to understand why I'm using a certain dispersion method and not another in DEseq (not DEseq2). Is there a way to get specific math and statistic skills to understand R properly?

    Leave a comment:


  • fpesce
    replied
    Originally posted by Michael Love View Post
    fpesce,

    sorry I can't replicate any bug with mixed numeric and categorical variables for >100 samples. If you can email me a small reproducible example (maybe subsetting the rows of dds) I can look into it. My email is listed in help(package="DESeq2")

    But I would recommend turning age into a categorical variable. You can split it into biologically meaningful groups: <18 y/o, 19-35, etc. Then the average effect of each group is controlled for, regardless of the trend over years. This often makes more sense rather than assuming a constant log2 fold change per year.

    You can turn numerics into factors using cut():

    > age <- 40:50
    > cut(age, breaks=c(0,30,45,60,200))
    [1] (30,45] (30,45] (30,45] (30,45] (30,45] (30,45] (45,60] (45,60] (45,60]
    [10] (45,60] (45,60]
    Levels: (0,30] (30,45] (45,60] (60,200]
    Thank you so much Mike, this completely makes sense and actually solved this issue.

    Leave a comment:


  • Michael Love
    replied
    IsBeth,

    There is information about the methods used by estimateDispersions in the ?estimateDispersions man page.

    In DESeq2, the Cox-Reid estimator (a penalized maximum likelihood estimator) is the only method provided, as it covers all experimental designs.

    it is also discussed in this man page what happens in the case of a 1 vs 1 comparison.

    Leave a comment:


  • Michael Love
    replied
    fpesce,

    sorry I can't replicate any bug with mixed numeric and categorical variables for >100 samples. If you can email me a small reproducible example (maybe subsetting the rows of dds) I can look into it. My email is listed in help(package="DESeq2")

    But I would recommend turning age into a categorical variable. You can split it into biologically meaningful groups: <18 y/o, 19-35, etc. Then the average effect of each group is controlled for, regardless of the trend over years. This often makes more sense rather than assuming a constant log2 fold change per year.

    You can turn numerics into factors using cut():

    > age <- 40:50
    > cut(age, breaks=c(0,30,45,60,200))
    [1] (30,45] (30,45] (30,45] (30,45] (30,45] (30,45] (45,60] (45,60] (45,60]
    [10] (45,60] (45,60]
    Levels: (0,30] (30,45] (45,60] (60,200]

    Leave a comment:


  • IsBeth
    replied
    Estimate Dispersions with DESeq2

    Hello! =)

    I'm a beginner (Blutiger Anfänger) at bioinformatics and so I've got a beginners question, sorry! Nowadays I'm trying to run DESeq2 with a dataset without replicates. It's a dataset with two columns (2 conditions, "control" and "infected").
    Which method do you recommend me to estimate the dispersion of my data? ( "pooled", "pooled-CR", "per-condition", "blind"?). I was reading about these methods but it's difficult for me to say which one I should use.

    If I just use the command "estimateDispersions" without specifying the method to make use of, which one is applied?

    What should I be careful with?


    Thanks in advance!

    Leave a comment:


  • fpesce
    replied
    Originally posted by Michael Love View Post
    Can you please post colData(dds) (or some scrubbed version of it)? This helps me figure out what the design matrix looks like, and therefore why some designs are causing errors.
    Here it is:
    > colData(dds)
    DataFrame with 122 rows and 4 columns
    CONDITION KEEP AGE SEX
    <factor> <integer> <numeric> <factor>
    ID1 A 1 47.06913 F
    ID2 A 1 45.33333 M
    ID3 A 1 59.63039 F
    ID4 A 1 49.00753 M
    ID5 A 1 34.94319 M
    ... ... ... ... ...
    ID118 B 1 30.00684 M
    ID119 B 1 41.90828 F
    ID120 B 1 39.04449 F
    ID121 B 1 39.64682 M
    ID122 B 1 16.70363 M

    Leave a comment:


  • Michael Love
    replied
    Can you please post colData(dds) (or some scrubbed version of it)? This helps me figure out what the design matrix looks like, and therefore why some designs are causing errors.

    Leave a comment:


  • fpesce
    replied
    Originally posted by dpryan View Post
    Could "AGE" have been converted to a factor at some point?
    No it has not. I checked:

    > is(colData(dds)$AGE)
    [1] "numeric" "vector" "atomic" "vectorORfactor"

    Thanks for the suggestion

    Leave a comment:


  • dpryan
    replied
    Could "AGE" have been converted to a factor at some point?

    Leave a comment:


  • fpesce
    replied
    Hi,
    I am using DESeq2 version 1.2.5 and when I use this design:

    > design(dds) <- formula(~ AGE + SEX + CONDITION)

    Defining the variables:
    > colData(dds)$CONDITION <- relevel(colData(dds)$CONDITION, "R")
    > colData(dds)$SEX <- relevel(colData(dds)$SEX, "M")
    Whereas AGE is a continuous variable

    I get this error:
    > dds <- DESeq(dds)
    using pre-existing size factors
    estimating dispersions
    you had estimated dispersions, replacing these
    gene-wise dispersion estimates
    error: inv(): matrix appears to be singular

    On the other hand, there are no problems with these combinations:
    > design(dds) <- formula(~ AGE)
    > design(dds) <- formula(~ CONDITION)
    > design(dds) <- formula(~ SEX + CONDITION)

    While I get the same error with:
    > design(dds) <- formula(~ AGE + CONDITION)

    Any clue what's going on here?
    Thank you so much in advance for your help

    Leave a comment:


  • David [R]
    replied
    contrast argument

    Hi,

    I have started using DESeq2 on a multi-factor design. Everything goes fine until I try to use the "contrast" argument on the "results" function in order to get somparisons other than the ones concerning the control condition.

    This is what I get:
    > res.R24.OE1=results(dds.gtype,contrast=c("gtype","R24","OE1"))
    Error in results(dds.gtype, contrast = c("gtype", "R24", "OE1")) :
    unused argument (contrast = c("gtype", "R24", "OE1"))

    "gtype" being my factor and "R24" and "OE1" my two conditions.

    I would be very grateful if you could help me out on this one.

    Thanks!

    Leave a comment:


  • sindrle
    replied
    Can I ask why you set it to .1 ?

    Leave a comment:


  • Simon Anders
    replied
    Please use
    Code:
    resSig <- res[ which( res$padj < .1 ), ]
    to get rid of the NAs.

    Leave a comment:


  • alisrpp
    replied
    Hi,

    I'm using DESeq2_1.2.0 and I am having this error:

    Code:
    > resSig <- res[ res$padj < .1, ]
    Error in normalizeSingleBracketSubscript(i, x, byrow = TRUE, exact = FALSE) : 
      subscript contains NAs
    It's really weird because this error started today and I have been using the same datasets and the same Rscript for the last 3 weeks without this error coming up.

    Just in case, here is my code:

    Code:
    #Count matrix input
    Cele_SPvsLR_old = read.csv (file.choose(), header=TRUE, row.names=1)
    CeleDesign <- data.frame(
      row.names = colnames(Cele_SPvsLR_old),
      condition = factor(c("SP", "SP", "LR", "LR")))
    dds <- DESeqDataSetFromMatrix(countData = Cele_SPvsLR_old,
                                  colData = CeleDesign,
                                  design = ~ condition)
    dds
    
    #Est size factor = normalize for library size
    dds<- estimateSizeFactors(dds)
    ddsLocal <- estimateDispersions(dds, fitType="local")
    ddsLocal <- nbinomWaldTest(ddsLocal)
    plotDispEsts(ddsLocal)
    
    #Differential expression analysis
    resultsNames(ddsLocal)
    res <- results(ddsLocal, name= "condition_SP_vs_LR") 
    res <- res[order(res$padj),]
    head(res)
    plotMA(ddsLocal, ylim=c(-2,2), main="DESeq2")
    sum(res$padj < .1, na.rm=TRUE)
    
    #filter for significant genes 
    resSig <- res[ res$padj < 0.1, ]
    And my sessionInfo()

    Code:
    > sessionInfo()
    R version 3.0.2 (2013-09-25)
    Platform: x86_64-apple-darwin10.8.0 (64-bit)
    
    locale:
    [1] en_US.UTF-8/en_US.UTF-8/en_US.UTF-8/C/en_US.UTF-8/en_US.UTF-8
    
    attached base packages:
    [1] parallel  stats     graphics  grDevices utils     datasets  methods   base     
    
    other attached packages:
    [1] DESeq2_1.2.0            RcppArmadillo_0.3.920.1 Rcpp_0.10.5             GenomicRanges_1.14.1   
    [5] XVector_0.2.0           IRanges_1.20.0          BiocGenerics_0.8.0     
    
    loaded via a namespace (and not attached):
     [1] annotate_1.40.0      AnnotationDbi_1.24.0 Biobase_2.22.0       DBI_0.2-7           
     [5] genefilter_1.44.0    grid_3.0.2           lattice_0.20-24      locfit_1.5-9.1      
     [9] RColorBrewer_1.0-5   RSQLite_0.11.4       splines_3.0.2        stats4_3.0.2        
    [13] survival_2.37-4      tools_3.0.2          XML_3.95-0.2         xtable_1.7-1
    Can anyone give me an explanation about why am I having this error now but not in the past using exactly the same Rscript and datasets? Does anyone know how to fix it?

    Thanks!
    Last edited by alisrpp; 10-24-2013, 11:50 AM.

    Leave a comment:


  • i4u412
    replied
    Hi

    I just started using DEseq2, so please correct if i am wrong.
    My sample description has five time points (with 2 replicates).
    (t0,t0,t1,t1,t2,t2,t3,t3,t4,t4,t5,t5)

    if i set the timepoint design, i always get the differential expression through initial time point (control).

    In addition to this, i also want to have differential expression with respect to the immediate time points (like t1 to t0, t2 to t1, t3 to t2, t4 to t3 and t5 to t4). Can you please tell me how to set the factor levels to get this type of comparison.

    thank you
    Last edited by i4u412; 10-23-2013, 04:57 AM.

    Leave a comment:

Latest Articles

Collapse

  • SEQadmin2
    New Genomics Technologies Take Aim at Long-Standing Limits
    by SEQadmin2


    Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

    We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
    ...
    09-28-2026, 10:25 AM

ad_right_rmr

Collapse

News

Collapse

Topics Statistics Last Post
Started by SEQadmin2, 09-29-2026, 09:51 AM
0 responses
43 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-25-2026, 09:06 AM
0 responses
53 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-23-2026, 11:05 AM
0 responses
43 views
0 reactions
Last Post SEQadmin2  
Started by SEQadmin2, 09-18-2026, 11:37 AM
1 response
51 views
0 reactions
Last Post pekgio
by pekgio
 
Working...