Originally posted by westerman
View Post
Seqanswers Leaderboard Ad
Collapse
Announcement
Collapse
No announcement yet.
X
-
-
Originally posted by KevinLam View PostI can't understand why they can't put the timestamp in the filename instead of the directory as well.
such a pain
primary.date1/reads/run_F3_sample.csfasta
primary.date2/reads/run_R3_sample.csfasta
Or however you have your experiment set up. The point is that the file names (not the directories) look similar thus making it easy to switch files in and out of the pipeline. Whereas with file stamps in the file names it would look like:
primary.date1/reads/run_F3_sample_date1.csfasta
primary.date2/reads/run_R3_sample_date2.csfasta
Not much of a change but perhaps something that Bioscope does not want to handle. Although, granted, making the individual files a lot more trackable.
Anyway the above is just being "the devil's advocate". Getting back to the original question:
Our people are unable to tell me which primary.* files to use and the 'latest' primary.* file doesn't always have the 'latest' write-to time stamp (unix time attribute) off the machine- nor is it the best quality.
In our (Purdue Genomics) case the only time we do a reanalysis mid-run is when something is screwed up -- i.e., we are desperately trying to save a run from a total meltdown. Fortunately the SOLiD technology is really good at being able redo parts of a run and extract data from a potentially failed run. However there have been cases where we redid a partial part of run, did not like that, re-did that part again, and ended up with even worse results. At that point we had a jumble of primary.* directories from which to choose the best of many poor results. You may be in a similar situation -- i.e., none of the results will be that great. But one of "your people" should be able to say, via looking at the heat maps and the cycle scans, which primary.* has the best chance of containing good data.
I know that the above doesn't really answer your question "... [a] way to ID the best-run and best primary* file to use ..." but this is in part because any time a reanalysis is done then implies that there is likely to be no "best". You have a non-typical case and thus are charting a non-typical path.
Leave a comment:
-
Sorry not quite getting you.
But if you mean how to set a rule to find the proper F3.csfasta files to use.
then yes I know wat you mean.
I usually do a
find . -iname *.stats
only the correct reads dir with a .stats file contains the F3 that I need.
I can't understand why they can't put the timestamp in the filename instead of the directory as well.
such a pain
Leave a comment:
-
Thanks, mrawlins. Maybe someone can also shed some light on the primary* issue...
Leave a comment:
-
If you disable the auto-export and export them manually using scp you can rename the files when you copy them. This doesn't solve the problem of finding which primary.* to use, though in our case it's always been the one with the largest number (most recent timestamp).
Leave a comment:
-
Finding optimal reads generated on SOLiD:
Hi,
We're having some issues with the data coming off the SOLiD. Our people are unable to tell me which primary.* files to use and the 'latest' primary.* file doesn't always have the 'latest' write-to time stamp (unix time attribute) off the machine- nor is it the best quality. Similarly, the primary.* directories may differ across the run from sample to sample and so one sample might use primary.20101016etc on one sample and another sample of the same exp/lib/run might use primary.20101014 and both are present in each directory. These are all attributable to doing reanalysis mid-seq during a run.
To make matters worse, the way the files are being copied is creating an entire system/network of directories that I have to go through to find the proper primary.* file for mapping. Using the wrong one results in very poor mapping. Could someone shed some light on a manual to look through or a way to ID the best-run and best primary* file to use off a run? Finally, is there a way to name these directories prior to export off the SOLiD? Thanks.
JTags: None
Latest Articles
Collapse
-
by seqadmin
The human gut contains trillions of microorganisms that impact digestion, immune functions, and overall health1. Despite major breakthroughs, we’re only beginning to understand the full extent of the microbiome’s influence on health and disease. Advances in next-generation sequencing and spatial biology have opened new windows into this complex environment, yet many questions remain. This article highlights two recent studies exploring how diet influences microbial...-
Channel: Articles
02-24-2025, 06:31 AM -
ad_right_rmr
Collapse
News
Collapse
Topics | Statistics | Last Post | ||
---|---|---|---|---|
Started by seqadmin, 03-03-2025, 01:15 PM
|
0 responses
171 views
0 likes
|
Last Post
by seqadmin
03-03-2025, 01:15 PM
|
||
Started by seqadmin, 02-28-2025, 12:58 PM
|
0 responses
261 views
0 likes
|
Last Post
by seqadmin
02-28-2025, 12:58 PM
|
||
Started by seqadmin, 02-24-2025, 02:48 PM
|
0 responses
644 views
0 likes
|
Last Post
by seqadmin
02-24-2025, 02:48 PM
|
||
Started by seqadmin, 02-21-2025, 02:46 PM
|
0 responses
265 views
0 likes
|
Last Post
by seqadmin
02-21-2025, 02:46 PM
|
Leave a comment: