Skip to content

Unexplainable mapping output when using two genome assemblies in rice plant #2688

Description

@vpbio

Hi Alex,
I will try to be as thourough as possible. In rice the most popular genome assembly is the MSU7 assembly, but there is a newer T2T assembly of AGIS1.0 also which is the current ref genome in NCBI.
I mapped 10-12 samples fom 10xGenomics 3' V3.1 libraries on AGIS assembly and got sample with maximum unique mapping at 90% and minimum unique read mapped sample with 42%. this is the final.log.out from the worst sample..

Started job on | Jun 06 02:15:10
Started mapping on | Jun 06 02:15:59
Finished on | Jun 06 03:34:24
Mapping speed, Million of reads per hour | 635.25
Number of input reads | 830240949
Average input read length | 77
UNIQUE READS:
Uniquely mapped reads number | 350381002
Uniquely mapped reads % | 42.20%
Average mapped length | 84.38
Number of splices: Total | 12779457
Number of splices: Annotated (sjdb) | 10454779
Number of splices: GT/AG | 11598524
Number of splices: GC/AG | 228976
Number of splices: AT/AC | 9364
Number of splices: Non-canonical | 942593
Mismatch rate per base, % | 0.23%
Deletion rate per base | 0.02%
Deletion average length | 1.59
Insertion rate per base | 0.02%
Insertion average length | 1.33
MULTI-MAPPING READS:
Number of reads mapped to multiple loci | 39865016
% of reads mapped to multiple loci | 4.80%
Number of reads mapped to too many loci | 1660559
% of reads mapped to too many loci | 0.20%
UNMAPPED READS:
Number of reads unmapped: too many mismatches | 0
% of reads unmapped: too many mismatches | 0.00%
Number of reads unmapped: too short | 138663071
% of reads unmapped: too short | 16.70%
Number of reads unmapped: other | 299671301
% of reads unmapped: other | 36.09%
CHIMERIC READS:
Number of chimeric reads | 0
% of chimeric reads | 0.00%

On this sample, to investigate further, I mapped with same parameters (mostly default) to MSU7 rice genome build and got ...

Started job on |	Jun 06 00:59:50
                         Started mapping on |	Jun 06 01:00:11
                                Finished on |	Jun 06 02:14:59
   Mapping speed, Million of reads per hour |	665.97

                      Number of input reads |	830240949
                  Average input read length |	77
                                UNIQUE READS:
               Uniquely mapped reads number |	349273174
                    Uniquely mapped reads % |	42.07%
                      Average mapped length |	84.37
                   Number of splices: Total |	13174974
        Number of splices: Annotated (sjdb) |	10427032
                   Number of splices: GT/AG |	11715549
                   Number of splices: GC/AG |	239872
                   Number of splices: AT/AC |	9140
           Number of splices: Non-canonical |	1210413
                  Mismatch rate per base, % |	0.22%
                     Deletion rate per base |	0.02%
                    Deletion average length |	1.60
                    Insertion rate per base |	0.02%
                   Insertion average length |	1.33
                         MULTI-MAPPING READS:
    Number of reads mapped to multiple loci |	400406907
         % of reads mapped to multiple loci |	48.23%
    Number of reads mapped to too many loci |	2180286
         % of reads mapped to too many loci |	0.26%
                              UNMAPPED READS:

Number of reads unmapped: too many mismatches | 0
% of reads unmapped: too many mismatches | 0.00%
Number of reads unmapped: too short | 74523151
% of reads unmapped: too short | 8.98%
Number of reads unmapped: other | 3857431
% of reads unmapped: other | 0.46%
CHIMERIC READS:
Number of chimeric reads | 0
% of chimeric reads | 0.00%

The most obvious difference is the significant no. of reads shifting from unmapped : other category to the reads mapped to multiple loci category. I went through the issues section to find out others having this issue and understood that it may be due to either contamination OR highly repeatitive sequences for which seeds matching to >50 (default) are reported in unmapped : other category.

So I followed the recommended approach of BLAST from unmapped reads (n=2000) and found that ~80% reads matches with rRNA sequences. Now, my hypothesis is that if the unmapped reads are from repeatitive elements (rRNA genes) I should be able to use the latest AGIS genome assembly and with relaxed parameters, I could get the reads mapped to multiple loci and try to increase some of the counts using STARsolo's expectation maximization algorithm. with following parameters (in addition to default usual parameters ---

--outFilterMultimapNmax 1000
--outFilterMatchNminOverLread 0.3
--winAnchorMultimapNmax 500
--alignWindowsPerReadNmax 50000
--alignTranscriptsPerWindowNmax 500
--seedMultimapNmax 50000
--seedPerWindowNmax 500
--seedPerReadNmax 5000
--seedNoneLociPerWindow 100

I am getting this output ---

       Time    Speed        Read     Read   Mapped   Mapped   Mapped   Mapped Unmapped Unmapped Unmapped Unmapped
                M/hr      number   length   unique   length   MMrate    multi   multi+       MM    short    other

Jun 08 17:53:13 116.7 2138643 77 41.8% 84.4 0.2% 5.4% 0.0% 0.0% 16.7% 36.0%
Jun 08 17:54:13 139.6 4884963 77 41.8% 84.4 0.2% 5.5% 0.0% 0.0% 16.6% 36.2%
Jun 08 17:55:14 145.6 7562625 77 41.8% 84.4 0.2% 5.5% 0.0% 0.0% 16.5% 36.2%
Jun 08 17:56:16 149.0 10306124 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.2%
Jun 08 17:57:16 151.0 12959333 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.2%
Jun 08 17:58:16 151.0 15476480 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 17:59:18 152.0 18197720 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 18:00:19 153.1 20918960 77 41.9% 84.4 0.2% 5.5% 0.0% 0.0% 16.4% 36.3%
Jun 08 18:01:19 154.2 23640200 77 41.9% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:02:19 154.7 26293409 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:03:22 154.4 28946618 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:04:23 155.6 31803920 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:05:25 156.1 34593191 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:06:25 156.6 37314431 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:07:25 156.7 39967640 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.3% 36.3%
Jun 08 18:08:25 156.6 42552818 77 42.0% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:09:28 156.8 45342089 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:10:28 157.4 48131360 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:11:29 157.5 50852600 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:12:29 157.6 53505809 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%
Jun 08 18:13:29 157.9 56227049 77 42.1% 84.4 0.2% 5.4% 0.0% 0.0% 16.2% 36.3%

My question is what am I missing here? Considering the reads multimap to msu7 genome I should be able to bypass the filters using AGIS genome and get the same output as MSU7 genome. Am I wrong in thinking that ? I already ruled out the genome differences because other samples gives good (90%) mapping with AGIS genome under default parameters. I cannot wrap my head around that how a T2T genome could give output where I cannot find what is happening. Let me know if additional info is needed.

Thanks,
Vivek

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions