Hello! This is my first post to SeqAnswers, but I have perused the forums before, so before I start describing my problem, hello to this fantastic community.
I'm hoping for some guidance on pooling libraries for sequencing on a NovaSeq X lane. Our group has 22 libraries that we want to pool on one NovaSeq X 10B lane. The trouble is, we need to pool the libraries in unequal molar ratios, because some libraries are being deep-sequenced, whilst others are being screened and we therefore only need low-depth sequencing <1 million clusters or Paired End reads. The proportions of the libraries are as follows:
When we pool this number of samples (20+), we wouldn't normally worry too much about the index balance. But the fact that the pool is dominated by four libraries, making up around ~95% of the pool, changes the balance of bases. Unfortunately, this didn't occur to me until after the pool had been made.
The four libraries that make up the pools have the following index combinations (our dual indexes are 6 bases each in length):
When you take into account the different representation of each library in the pool, we get a proportion of A/C and G/T bases as follows:
As you can see, there are some positions in both the P5 and P7 indexes that have very unequal proportions of A/C and G/T. I think if we use the spectral balance rules for a NovaSeq X (A/C in Blue Channel, C/T in Green Channel at every cycle, then I think the situation looks better, but I am still concerned as I have no experience with the NovaSeq X system. What impact might that have on sequencing the indexes and is there any way to mitigate this in a sequencing lane, such as increasing the percentage of phiX? I know that phiX doesn't have indexes, but I thought that more phiX might help with base balance during the calculation of spectral correction tables. My options are either to try to mitigate the problems with phiX or pull the pool from the sequencing queue and re-amplify the libraries with different indexes. It's unfortunate that we have to sequence this on a single lane, but that is one of the constraints that I am working with. Any estimates of how the base imbalance in the indexes might affect accurate basecalling of the indexes, such as whether we might expect to lose 10 or 20% of our index reads (and therefore sequencing reads), as well as advice on how to improve results without re-doing the pool entirely, would be very helpful. Thanks in advance!
I'm hoping for some guidance on pooling libraries for sequencing on a NovaSeq X lane. Our group has 22 libraries that we want to pool on one NovaSeq X 10B lane. The trouble is, we need to pool the libraries in unequal molar ratios, because some libraries are being deep-sequenced, whilst others are being screened and we therefore only need low-depth sequencing <1 million clusters or Paired End reads. The proportions of the libraries are as follows:
| Lab Code | Proportion |
| L1 | 0.28 |
| L2 | 0.38 |
| L3 | 0.0009 |
| L4 | 0.0009 |
| L5 | 0.0009 |
| L6 | 0.0009 |
| L7 | 0.0009 |
| L8 | 0.0009 |
| L9 | 0.005 |
| L10 | 0.005 |
| L11 | 0.005 |
| L12 | 0.005 |
| L13 | 0.005 |
| L14 | 0.005 |
| L15 | 0.005 |
| L16 | 0.005 |
| L17 | 0.005 |
| L18 | 0.005 |
| L19 | 0.005 |
| L20 | 0.14 |
| L21 | 0.14 |
| L22 | 0.0009 |
When we pool this number of samples (20+), we wouldn't normally worry too much about the index balance. But the fact that the pool is dominated by four libraries, making up around ~95% of the pool, changes the balance of bases. Unfortunately, this didn't occur to me until after the pool had been made.
The four libraries that make up the pools have the following index combinations (our dual indexes are 6 bases each in length):
| P5 | P7 |
| GCTATA | CGTATA |
| CGTCAG | CGCTAT |
| CTAGAT | GCAACG |
| CGACGT | TTAGGC |
When you take into account the different representation of each library in the pool, we get a proportion of A/C and G/T bases as follows:
| A/C | G/T | |
| P5-1 | 0.69 | 0.31 |
| P5-2 | 0.31 | 0.69 |
| P5-3 | 0.31 | 0.69 |
| P5-4 | 0.83 | 0.17 |
| P5-5 | 0.55 | 0.45 |
| P5-6 | 0.31 | 0.69 |
| P7-1 | 0.68 | 0.32 |
| P7-2 | 0.17 | 0.83 |
| P7-3 | 0.70 | 0.30 |
| P7-4 | 0.45 | 0.55 |
| P7-5 | 0.55 | 0.45 |
| P7-6 | 0.44 | 0.56 |
As you can see, there are some positions in both the P5 and P7 indexes that have very unequal proportions of A/C and G/T. I think if we use the spectral balance rules for a NovaSeq X (A/C in Blue Channel, C/T in Green Channel at every cycle, then I think the situation looks better, but I am still concerned as I have no experience with the NovaSeq X system. What impact might that have on sequencing the indexes and is there any way to mitigate this in a sequencing lane, such as increasing the percentage of phiX? I know that phiX doesn't have indexes, but I thought that more phiX might help with base balance during the calculation of spectral correction tables. My options are either to try to mitigate the problems with phiX or pull the pool from the sequencing queue and re-amplify the libraries with different indexes. It's unfortunate that we have to sequence this on a single lane, but that is one of the constraints that I am working with. Any estimates of how the base imbalance in the indexes might affect accurate basecalling of the indexes, such as whether we might expect to lose 10 or 20% of our index reads (and therefore sequencing reads), as well as advice on how to improve results without re-doing the pool entirely, would be very helpful. Thanks in advance!