# Too many duplicated sequences OR unassigned because of mapping quality

**URL:** <https://help.galaxyproject.org/t/too-many-duplicated-sequences-or-unassigned-because-of-mapping-quality/11965>\
**Category:** Uncategorized\
**Tags:** troubleshooting, mapping, transcriptomics, tool-help, featurecounts\
**Created:** [March 19, 2024, 4:50pm UTC](https://help.galaxyproject.org/t/too-many-duplicated-sequences-or-unassigned-because-of-mapping-quality/11965 "2024-03-19T16:50:45Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![MelanieP](https://avatars.discourse-cdn.com/v4/letter/m/fbc32d/32.png) [@MelanieP](https://help.galaxyproject.org/u/MelanieP)\
**Post date:** [March 19, 2024, 4:50pm UTC](https://help.galaxyproject.org/t/too-many-duplicated-sequences-or-unassigned-because-of-mapping-quality/11965/1 "2024-03-19T16:50:45Z")

</div>

Hello,

I have problems in the very beginning of my RNA sequence analysis and it would be great if somebody could help.

After importing the sequence data, I did cutadapt. The result is, that I have a rate of 88 % duplicates. Here is a picture of the Sequence Duplication Levels:  
 ![duplicates](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/2X/a/a62b7203f93a2e1852af29545bb5455407b05b6b.jpeg)

Then I did featureCounts and I got this:

![mapping quality](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/2X/1/10cc22f999134687b52b8a46d9f9ba06c56f32b2.jpeg)

I can’t find an explanation of the high rate of unassigned because of Multi Mapping.

After Counting and Annotation and so on, I figured out, that roughly 35% of my RNA belongs to ribosomal RNA.

Did anybody has experience with this and could help?

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [March 19, 2024, 11:47pm UTC](https://help.galaxyproject.org/t/too-many-duplicated-sequences-or-unassigned-because-of-mapping-quality/11965/2 "2024-03-19T23:47:08Z")

</div>

Welcome, @MelanieP

Multi-mapping is usually a problem with the annotation used during analysis, not the reads.

> [@MelanieP](#):
>
> the high rate of unassigned because of Multi Mapping

This was a problem introduced into the library construction _prior_ to sequencing. Those reads will probably fall out during mapping steps in a standard transcriptomics analysis protocol.

> [@MelanieP](#):
>
> ut, that roughly 35% of my RNA belongs to ribosomal RNA

More help for the reference annotation part, and potentially analysis methods →

- [FAQ: Extended Help for Differential Expression Analysis Tools](https://training.galaxyproject.org/training-material/faqs/galaxy/analysis_differential_expression_help.html)
- [Transcriptomics / Tutorial List](https://training.galaxyproject.org/training-material/topics/transcriptomics/)

For what to do about the ribosomal RNA outside of attempting to remove it during mapping, you should review scientific forums, publications, and related resources. You may have to decide whether the reads are usable at all.

---

<div class="post-metadata">

**Author:** ![pavanvidem](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/pavanvidem/32/1687_2.png) [@pavanvidem](https://help.galaxyproject.org/u/pavanvidem)\
**Post date:** [March 20, 2024, 1:29pm UTC](https://help.galaxyproject.org/t/too-many-duplicated-sequences-or-unassigned-because-of-mapping-quality/11965/3 "2024-03-20T13:29:50Z")

</div>

I would like to discuss a bit more about choosing the annotation, and strange sources of multi-mapping because I investigated this particular analysis. As @jennaj mentioned, the annotation can influence the `featureCounts`.

- If you have a decent number of uniquely mapped reads (in `RNA STAR` log) but many **Unassigned No feature** reads in `featureCounts` output, then try using a different annotation file. There is a good chance that the reads are from ribosomal RNAs (rRNA) but the GTF file does not have a complete annotation of rRNAs. Sometimes, you have a better annotation of rRNAs in Refseq annotation than Ensembl or Gencode annotations.

- In RNA-seq, you can expect ~50-60% duplication because of the reads coming from exons shared among isoforms. In your samples, the abnormally high duplication level is most likely from the high % of rRNAs.

- There is a relationship between the multi-mapped reads and their mapping quality. Alignment programs assign a low mapping quality (_MAPQ_ field in BAM files) to multi-mapped reads and a high value for uniquely mapped. These values depend on the aligner used. By default, `RNA STAR` uses a _MAPQ_ value of `60` to indicate a uniquely mapped read, and a _MAPQ_ of `3` to indicate a read multi-mapped to two genomic loci. This value is further decreased based on the number of hits on the reference genome. By default, in `featureCounts`, the **Minimum mapping quality per read** parameter is set to `0`. Hence, every multi-mapped alignment counted as **Unassigned\_MultiMapping**. If you set this parameter value to `10`, all the alignments with `MAPQ<10` are considered as **Unassigned\_MappingQuality** in the `featureCounts` output. So, depending on the value set for **Minimum mapping quality per read** parameter, you see them as either multi-mapped or low-quality alignments which are actually the same set of multi-mapped reads.

- Now the question is what causes multimapped reads and can we do something about it?
