Hi Alex,
I will try to be as thourough as possible. In rice the most popular genome assembly is the MSU7 assembly, but there is a newer T2T assembly of AGIS1.0 also which is the current ref genome in NCBI.
I mapped 10-12 samples fom 10xGenomics 3' V3.1 libraries on AGIS assembly and got sample with maximum unique mapping at 90% and minimum unique read mapped sample with 42%. this is the final.log.out from the worst sample..
Started job on | Jun 06 02:15:10
Started mapping on | Jun 06 02:15:59
Finished on | Jun 06 03:34:24
Mapping speed, Million of reads per hour | 635.25
Number of input reads | 830240949
Average input read length | 77
UNIQUE READS:
Uniquely mapped reads number | 350381002
Uniquely mapped reads % | 42.20%
Average mapped length | 84.38
Number of splices: Total | 12779457
Number of splices: Annotated (sjdb) | 10454779
Number of splices: GT/AG | 11598524
Number of splices: GC/AG | 228976
Number of splices: AT/AC | 9364
Number of splices: Non-canonical | 942593
Mismatch rate per base, % | 0.23%
Deletion rate per base | 0.02%
Deletion average length | 1.59
Insertion rate per base | 0.02%
Insertion average length | 1.33
MULTI-MAPPING READS:
Number of reads mapped to multiple loci | 39865016
% of reads mapped to multiple loci | 4.80%
Number of reads mapped to too many loci | 1660559
% of reads mapped to too many loci | 0.20%
UNMAPPED READS:
Number of reads unmapped: too many mismatches | 0
% of reads unmapped: too many mismatches | 0.00%
Number of reads unmapped: too short | 138663071
% of reads unmapped: too short | 16.70%
Number of reads unmapped: other | 299671301
% of reads unmapped: other | 36.09%
CHIMERIC READS:
Number of chimeric reads | 0
% of chimeric reads | 0.00%
On this sample, to investigate further, I mapped with same parameters (mostly default) to MSU7 rice genome build and got ...
Started job on | Jun 06 00:59:50
Started mapping on | Jun 06 01:00:11
Finished on | Jun 06 02:14:59
Mapping speed, Million of reads per hour | 665.97
Number of input reads | 830240949
Average input read length | 77
UNIQUE READS:
Uniquely mapped reads number | 349273174
Uniquely mapped reads % | 42.07%
Average mapped length | 84.37
Number of splices: Total | 13174974
Number of splices: Annotated (sjdb) | 10427032
Number of splices: GT/AG | 11715549
Number of splices: GC/AG | 239872
Number of splices: AT/AC | 9140
Number of splices: Non-canonical | 1210413
Mismatch rate per base, % | 0.22%
Deletion rate per base | 0.02%
Deletion average length | 1.60
Insertion rate per base | 0.02%
Insertion average length | 1.33
MULTI-MAPPING READS:
Number of reads mapped to multiple loci | 400406907
% of reads mapped to multiple loci | 48.23%
Number of reads mapped to too many loci | 2180286
% of reads mapped to too many loci | 0.26%
UNMAPPED READS:
Number of reads unmapped: too many mismatches | 0
% of reads unmapped: too many mismatches | 0.00%
Number of reads unmapped: too short | 74523151
% of reads unmapped: too short | 8.98%
Number of reads unmapped: other | 3857431
% of reads unmapped: other | 0.46%
CHIMERIC READS:
Number of chimeric reads | 0
% of chimeric reads | 0.00%
The most obvious difference is the significant no. of reads shifting from unmapped : other category to the reads mapped to multiple loci category. I went through the issues section to find out others having this issue and understood that it may be due to either contamination OR highly repeatitive sequences for which seeds matching to >50 (default) are reported in unmapped : other category.
So I followed the recommended approach of BLAST from unmapped reads (n=2000) and found that ~80% reads matches with rRNA sequences. Now, my hypothesis is that if the unmapped reads are from repeatitive elements (rRNA genes) I should be able to use the latest AGIS genome assembly and with relaxed parameters, I could get the reads mapped to multiple loci and try to increase some of the counts using STARsolo's expectation maximization algorithm. with following parameters (in addition to default usual parameters ---
--outFilterMultimapNmax 1000
--outFilterMatchNminOverLread 0.3
--winAnchorMultimapNmax 500
--alignWindowsPerReadNmax 50000
--alignTranscriptsPerWindowNmax 500
--seedMultimapNmax 50000
--seedPerWindowNmax 500
--seedPerReadNmax 5000
--seedNoneLociPerWindow 100
I am getting this output ---
Time Speed Read Read Mapped Mapped Mapped Mapped Unmapped Unmapped Unmapped Unmapped
M/hr number length unique length MMrate multi multi+ MM short other
Jun 08 17:53:13 116.7 2138643 77 41.8% 84.4 0.2% 5.4% 0.0% 0.0% 16.7% 36.0%
Jun 08 17:54:13 139.6 4884963 77 41.8% 84.4 0.2% 5.5% 0.0% 0.0% 16.6% 36.2%
Jun 08 17:55:14 145.6 7562625 77 41.8% 84.4 0.2% 5.5% 0.0% 0.0% 16.5% 36.2%
Jun 08 17:56:16 149.0 10306124 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.2%
Jun 08 17:57:16 151.0 12959333 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.2%
Jun 08 17:58:16 151.0 15476480 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 17:59:18 152.0 18197720 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 18:00:19 153.1 20918960 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 18:01:19 154.2 23640200 77 41.9% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:02:19 154.7 26293409 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:03:22 154.4 28946618 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:04:23 155.6 31803920 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:05:25 156.1 34593191 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:06:25 156.6 37314431 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:07:25 156.7 39967640 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:08:25 156.6 42552818 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:09:28 156.8 45342089 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:10:28 157.4 48131360 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:11:29 157.5 50852600 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:12:29 157.6 53505809 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:13:29 157.9 56227049 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
My question is what am I missing here? Considering the reads multimap to msu7 genome I should be able to bypass the filters using AGIS genome and get the same output as MSU7 genome. Am I wrong in thinking that ? I already ruled out the genome differences because other samples gives good (90%) mapping with AGIS genome under default parameters. I cannot wrap my head around that how a T2T genome could give output where I cannot find what is happening. Let me know if additional info is needed.
Thanks,
Vivek
Hi Alex,
I will try to be as thourough as possible. In rice the most popular genome assembly is the MSU7 assembly, but there is a newer T2T assembly of AGIS1.0 also which is the current ref genome in NCBI.
I mapped 10-12 samples fom 10xGenomics 3' V3.1 libraries on AGIS assembly and got sample with maximum unique mapping at 90% and minimum unique read mapped sample with 42%. this is the final.log.out from the worst sample..
Started job on | Jun 06 02:15:10
Started mapping on | Jun 06 02:15:59
Finished on | Jun 06 03:34:24
Mapping speed, Million of reads per hour | 635.25
Number of input reads | 830240949
Average input read length | 77
UNIQUE READS:
Uniquely mapped reads number | 350381002
Uniquely mapped reads % | 42.20%
Average mapped length | 84.38
Number of splices: Total | 12779457
Number of splices: Annotated (sjdb) | 10454779
Number of splices: GT/AG | 11598524
Number of splices: GC/AG | 228976
Number of splices: AT/AC | 9364
Number of splices: Non-canonical | 942593
Mismatch rate per base, % | 0.23%
Deletion rate per base | 0.02%
Deletion average length | 1.59
Insertion rate per base | 0.02%
Insertion average length | 1.33
MULTI-MAPPING READS:
Number of reads mapped to multiple loci | 39865016
% of reads mapped to multiple loci | 4.80%
Number of reads mapped to too many loci | 1660559
% of reads mapped to too many loci | 0.20%
UNMAPPED READS:
Number of reads unmapped: too many mismatches | 0
% of reads unmapped: too many mismatches | 0.00%
Number of reads unmapped: too short | 138663071
% of reads unmapped: too short | 16.70%
Number of reads unmapped: other | 299671301
% of reads unmapped: other | 36.09%
CHIMERIC READS:
Number of chimeric reads | 0
% of chimeric reads | 0.00%
On this sample, to investigate further, I mapped with same parameters (mostly default) to MSU7 rice genome build and got ...
Number of reads unmapped: too many mismatches | 0
% of reads unmapped: too many mismatches | 0.00%
Number of reads unmapped: too short | 74523151
% of reads unmapped: too short | 8.98%
Number of reads unmapped: other | 3857431
% of reads unmapped: other | 0.46%
CHIMERIC READS:
Number of chimeric reads | 0
% of chimeric reads | 0.00%
The most obvious difference is the significant no. of reads shifting from unmapped : other category to the reads mapped to multiple loci category. I went through the issues section to find out others having this issue and understood that it may be due to either contamination OR highly repeatitive sequences for which seeds matching to >50 (default) are reported in unmapped : other category.
So I followed the recommended approach of BLAST from unmapped reads (n=2000) and found that ~80% reads matches with rRNA sequences. Now, my hypothesis is that if the unmapped reads are from repeatitive elements (rRNA genes) I should be able to use the latest AGIS genome assembly and with relaxed parameters, I could get the reads mapped to multiple loci and try to increase some of the counts using STARsolo's expectation maximization algorithm. with following parameters (in addition to default usual parameters ---
--outFilterMultimapNmax 1000
--outFilterMatchNminOverLread 0.3
--winAnchorMultimapNmax 500
--alignWindowsPerReadNmax 50000
--alignTranscriptsPerWindowNmax 500
--seedMultimapNmax 50000
--seedPerWindowNmax 500
--seedPerReadNmax 5000
--seedNoneLociPerWindow 100
I am getting this output ---
Jun 08 17:53:13 116.7 2138643 77 41.8% 84.4 0.2% 5.4% 0.0% 0.0% 16.7% 36.0%
Jun 08 17:54:13 139.6 4884963 77 41.8% 84.4 0.2% 5.5% 0.0% 0.0% 16.6% 36.2%
Jun 08 17:55:14 145.6 7562625 77 41.8% 84.4 0.2% 5.5% 0.0% 0.0% 16.5% 36.2%
Jun 08 17:56:16 149.0 10306124 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.2%
Jun 08 17:57:16 151.0 12959333 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.2%
Jun 08 17:58:16 151.0 15476480 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 17:59:18 152.0 18197720 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 18:00:19 153.1 20918960 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 18:01:19 154.2 23640200 77 41.9% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:02:19 154.7 26293409 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:03:22 154.4 28946618 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:04:23 155.6 31803920 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:05:25 156.1 34593191 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:06:25 156.6 37314431 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:07:25 156.7 39967640 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:08:25 156.6 42552818 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:09:28 156.8 45342089 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:10:28 157.4 48131360 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:11:29 157.5 50852600 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:12:29 157.6 53505809 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:13:29 157.9 56227049 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
My question is what am I missing here? Considering the reads multimap to msu7 genome I should be able to bypass the filters using AGIS genome and get the same output as MSU7 genome. Am I wrong in thinking that ? I already ruled out the genome differences because other samples gives good (90%) mapping with AGIS genome under default parameters. I cannot wrap my head around that how a T2T genome could give output where I cannot find what is happening. Let me know if additional info is needed.
Thanks,
Vivek