# Unassigned Multimapping in featurecounts

**URL:** <https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399>\
**Category:** Uncategorized\
**Tags:** reference-annotation, salmon\
**Created:** [October 31, 2019, 6:05pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399 "2019-10-31T18:05:07Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![srashid](https://avatars.discourse-cdn.com/v4/letter/s/ee7513/32.png) [@srashid](https://help.galaxyproject.org/u/srashid)\
**Post date:** [October 31, 2019, 6:05pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/1 "2019-10-31T18:05:08Z")

</div>

I have paired end fastq files from illumina Novaseq using whole transcriptome mRNA-seq profiling. My RNA STAR result looks OK (using hg38 gtf file from ucsc table browser).

```
                      Number of input reads |	38164847
                  Average input read length |	201
                                UNIQUE READS:
               Uniquely mapped reads number |	30007228
                    Uniquely mapped reads % |	78.63%
                      Average mapped length |	200.97
                   Number of splices: Total |	18124401
        Number of splices: Annotated (sjdb) |	17921546
                   Number of splices: GT/AG |	17970447
                   Number of splices: GC/AG |	117674
                   Number of splices: AT/AC |	16850
           Number of splices: Non-canonical |	19430
                  Mismatch rate per base, % |	0.19%
                     Deletion rate per base |	0.01%
                    Deletion average length |	1.73
                    Insertion rate per base |	0.01%
                   Insertion average length |	1.47
                         MULTI-MAPPING READS:
    Number of reads mapped to multiple loci |	7150744
         % of reads mapped to multiple loci |	18.74%
    Number of reads mapped to too many loci |	62501
         % of reads mapped to too many loci |	0.16%
                              UNMAPPED READS:
   % of reads unmapped: too many mismatches |	0.00%
             % of reads unmapped: too short |	2.44%
                 % of reads unmapped: other |	0.03%
                              CHIMERIC READS:
                   Number of chimeric reads |	0
                        % of chimeric reads |	0.00%

```

However, when I use featureCounts I get very few assigned reads, most of the reads are in Unassigned Multimapped category. (I use reverse stranded option as indicated by ‘infer experiment’)

| Assigned | 7712538 |
| --- | --- |
| Unassigned\_Unmapped | 0 |
| Unassigned\_MappingQuality | 0 |
| Unassigned\_Chimera | 0 |
| Unassigned\_FragmentLength | 0 |
| Unassigned\_Duplicate | 0 |
| Unassigned\_MultiMapping | 37553849 |
| Unassigned\_Secondary | 0 |
| Unassigned\_NonSplit | 0 |
| Unassigned\_NoFeatures | 4138912 |
| Unassigned\_Overlapping\_Length | 0 |
| Unassigned\_Ambiguity | 18155778 |

What could be the reason behind it? Is there a way to improve on that? I am losing a lot of reads due to multimapping.

I also checked the read distribution. It looks OK to me too.

 ![rseqc_read_distribution_plot](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/2X/b/b1d3e603023d5d32d3208ba04d797c72fae138c2.png)

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [October 31, 2019, 7:55pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/2 "2019-10-31T19:55:19Z")

</div>

Welcome @srashid

Avoid the UCSC reference GTFs from their Table Browser. These often end up truncated, plus there is a serious data content concern. Why is covered in this FAQ in more detail:

- [Extended Help for Differential Expression Analysis Tools](https://galaxyproject.org/support/diff-expression/)

Good sources for hg38 GTF reference annotation are described in this prior Q&A (and are included in the FAQ above as well):

> [@RNA-STAR and hg38 GTF reference annotation](https://help.galaxyproject.org/t/rna-star-and-hg38-gtf-reference-annotation/749/2):
>
> The GTF should be based on the UCSC “hg38” genome build. Some choices:
> 
> - For **Gencode** , copy the link to the GTF and paste it into the _Upload_ tool. Hg38 data is here [https://www.gencodegenes.org/](https://www.gencodegenes.org/). After it is loaded, remove the headers (lines that start with a “#”) with the _Select_ tool using the options “NOT Matching” with the regular expression `^#` . Once the formatting is fixed, change the datatype to be `gft` under Edit Attributes (pencil icon). The data will be given the datatype `gff` by default, which works fine with some tools and but not with others. Avoid the `gff3` version of this particular data (contains duplicated IDs and several RNA-seq tools do not work with annotation in that format anyway).
> - For **iGenomes** , the archive corresponding to the target genome/build needs to be locally downloaded, the tar archive unpacked, and then just the `genes.gtf` data uploaded to Galaxy (browse the local file, or use FTP). Find all available genome/builds here: [iGenomes](https://support.illumina.com/sequencing/sequencing_software/igenome.html)

Give one or both of those a try and see if your “Unassigned\_Ambiguity” and “Unassigned\_MultiMapping” counts reduce – they should (“gene\_id” and “transcript\_id” will no longer be the same value).

You may even get fewer “Unassigned\_NoFeatures” if the UCSC data was truncated when extracted from the Table Browser.

---

<div class="post-metadata">

**Author:** ![srashid](https://avatars.discourse-cdn.com/v4/letter/s/ee7513/32.png) [@srashid](https://help.galaxyproject.org/u/srashid)\
**Post date:** [October 31, 2019, 8:10pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/3 "2019-10-31T20:10:15Z")

</div>

Thank you! While waiting for your reply, I actually tried igenomes gtf file, it definitely reduces ambiguity , but the multimapping issue still remains

This is the output of featureCounts now

| Assigned | 23399833 |
| --- | --- |
| Unassigned\_Unmapped | 0 |
| Unassigned\_MappingQuality | 0 |
| Unassigned\_Chimera | 0 |
| Unassigned\_FragmentLength | 0 |
| Unassigned\_Duplicate | 0 |
| Unassigned\_MultiMapping | 37557436 |
| Unassigned\_Secondary | 0 |
| Unassigned\_NonSplit | 0 |
| Unassigned\_NoFeatures | 6348801 |
| Unassigned\_Overlapping\_Length | 0 |
| Unassigned\_Ambiguity | 259191 |

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [November 1, 2019, 4:19pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/4 "2019-11-01T16:19:52Z")

</div>

Hi @srashid

`FeatureCounts` only reports unique matches with default settings.

Your reads are likely hitting more than one “exon”, which leads to “multimapping” counts when summarized at the Gene level.

Review the “Advanced Options”. In particular, pay attention to these parameters, but also review others and see what results. There isn’t a single right answer for everyone. It depends on how you want these counted up, if at all.

- “Allow read to contribute to multiple features” (default=no)
- “Largest overlap” (default=no)
- “Count multi-mapping reads/fragments” (default=disabled) and the sub-option (when enabled) “Assign fractions to multimapping reads”

Thanks!

---

<div class="post-metadata">

**Author:** ![dartagnan32](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/dartagnan32/32/3083_2.png) [@dartagnan32](https://help.galaxyproject.org/u/dartagnan32)\
**Post date:** [March 24, 2021, 1:31pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/5 "2021-03-24T13:31:20Z")

</div>

Did you resolve your issue? it would be interesting to know what the solution was.

---

<div class="post-metadata">

**Author:** ![mmomeni](https://avatars.discourse-cdn.com/v4/letter/m/a587f6/32.png) [@mmomeni](https://help.galaxyproject.org/u/mmomeni)\
**Post date:** [June 11, 2021, 12:30pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/6 "2021-06-11T12:30:38Z")

</div>

Hi.  
I have the same problem. A lot of unassigned Multimapped reads. When I select the `Allow reads to map to multiple features` option, My problem is fixed and more than 50% of reads are assigned. Now I wanted to ask is it scientifically and technically okay to allow reads to map to multiple features? (My aim from this analysis is to find DEGs)

---

<div class="post-metadata">

**Author:** ![David](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/david/32/7031_2.png) [@David](https://help.galaxyproject.org/u/David)\
**Post date:** [June 11, 2021, 2:41pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/7 "2021-06-11T14:41:35Z")

</div>

@mmomeni, you can find some discussions/methods here:

[https://www.biostars.org/p/273609/](https://www.biostars.org/p/273609/)

[https://doi.org/10.1016/j.csbj.2020.06.014](https://doi.org/10.1016/j.csbj.2020.06.014)

[https://www.biostars.org/p/311322/](https://www.biostars.org/p/311322/)

---

<div class="post-metadata">

**Author:** ![mmomeni](https://avatars.discourse-cdn.com/v4/letter/m/a587f6/32.png) [@mmomeni](https://help.galaxyproject.org/u/mmomeni)\
**Post date:** [June 13, 2021, 12:29pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/8 "2021-06-13T12:29:58Z")

</div>

Thanks for introducing these discussions. I read them.  
Can we use one of the tools Salmon, RSEM, or Kallisto in Galaxy for dealing with multi-mapped reads?  
If the answer is yes, does any tutorials exist for that? If Galaxy has any other tool for this aim please introduce to me.

---

<div class="post-metadata">

**Author:** ![gallardoalba](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/gallardoalba/32/1903_2.png) [@gallardoalba](https://help.galaxyproject.org/u/gallardoalba)\
**Post date:** [June 14, 2021, 7:02pm UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/9 "2021-06-14T19:02:48Z")

</div>

Hi @mmomeni,  
yes, [Kallisto](https://usegalaxy.eu/root?tool_id=toolshed.g2.bx.psu.edu/repos/iuc/kallisto_quant/kallisto_quant/0.46.0.4), [RSEM](https://usegalaxy.eu/root?tool_id=toolshed.g2.bx.psu.edu/repos/jjohnson/rsem/rsem_calculate_expression/1.1.17) and [Salmon](https://usegalaxy.eu/root?tool_id=toolshed.g2.bx.psu.edu/repos/bgruening/salmon/salmon/0.14.1.2+galaxy1) are available in Galaxy. I recommed you to have a look at this tutorial in order to learn how to use Salmon for gene quantification: [Quantification of gene expression: Salmon](https://training.galaxyproject.org/training-material/topics/transcriptomics/tutorials/mirna-target-finder/tutorial.html#quantification-of-gene-expression-salmon).

Regards

---

<div class="post-metadata">

**Author:** ![mmomeni](https://avatars.discourse-cdn.com/v4/letter/m/a587f6/32.png) [@mmomeni](https://help.galaxyproject.org/u/mmomeni)\
**Post date:** [June 15, 2021, 4:17am UTC](https://help.galaxyproject.org/t/unassigned-multimapping-in-featurecounts/2399/10 "2021-06-15T04:17:48Z")

</div>

Thanks a lot. that was interesting tutorial and interesting tool!
