Hullo! Apologies for the seemingly basic question but I am in knots over this and hope there is some help.
I have recently 'inherited' an Illumina dataset containing about 36,000SNP genotypes for about 400 individuals. The data I have is presented as a 3 columns in a text file.
The columns are 'Indiv_ID', 'SNP_ID' and 'Genotype' (so 36000x400 rows in total) and I need the data in some usable format so that I can extract data for specific 500 SNPs. ideally I would prefer the data in an individual x SNP matrix.
Usually I have used R to reshape such data as such but this file seems to be too big and it just freezes whilst processing. I have also used Plink in the past to extract specific SNP data from 'column' data but in the text file I have, the genotype is given as an 'AB' format which Plink doesn't accept as a compound genotype. I have attempted to change all As and Bs for 1s and 2s so that I can input as a compound to Plink, but the software I was using to do this also adds " "s which then need to be removed. And no text editor I have seems to cope with this for so many lines of text.
I am not able to get this data in any other format and the only other data I have in relation to this is (just) enough for me to creat a .map file for Plink. Otherwise it is just the (seeminly) infinite column. I appreciate that this is probably a very simple task all in but at the moment I cannot see the wood for trees, and am going around in circles. I would welcome any starting points or good reference sites to check out!
Thanks!
I have recently 'inherited' an Illumina dataset containing about 36,000SNP genotypes for about 400 individuals. The data I have is presented as a 3 columns in a text file.
The columns are 'Indiv_ID', 'SNP_ID' and 'Genotype' (so 36000x400 rows in total) and I need the data in some usable format so that I can extract data for specific 500 SNPs. ideally I would prefer the data in an individual x SNP matrix.
Usually I have used R to reshape such data as such but this file seems to be too big and it just freezes whilst processing. I have also used Plink in the past to extract specific SNP data from 'column' data but in the text file I have, the genotype is given as an 'AB' format which Plink doesn't accept as a compound genotype. I have attempted to change all As and Bs for 1s and 2s so that I can input as a compound to Plink, but the software I was using to do this also adds " "s which then need to be removed. And no text editor I have seems to cope with this for so many lines of text.
I am not able to get this data in any other format and the only other data I have in relation to this is (just) enough for me to creat a .map file for Plink. Otherwise it is just the (seeminly) infinite column. I appreciate that this is probably a very simple task all in but at the moment I cannot see the wood for trees, and am going around in circles. I would welcome any starting points or good reference sites to check out!
Thanks!
Comment