# Removing PCR duplicates in GStacks2, galaxy.eu

**URL:** <https://help.galaxyproject.org/t/removing-pcr-duplicates-in-gstacks2-galaxy-eu/16393>\
**Category:** Uncategorized\
**Tags:** tool-help, stacks2\_denovomap, stacks\_denovomap, stacks2\_ustacks, stacks\_ustacks\
**Created:** [October 15, 2025, 2:34pm UTC](https://help.galaxyproject.org/t/removing-pcr-duplicates-in-gstacks2-galaxy-eu/16393 "2025-10-15T14:34:12Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [October 18, 2025, 12:39am UTC](https://help.galaxyproject.org/t/removing-pcr-duplicates-in-gstacks2-galaxy-eu/16393/2 "2025-10-18T00:39:06Z")

</div>

Hi @Silvia

This is a good followup question, so I’m glad you asked!

The duplicate removal is possible with gstacks when there is a reference because the reads are first mapped to it (and is why the input is an aligned BAM, not fastq). This sets up a coordinate system where “exactly repeated” mapping characteristics can then be filtered out as PCR duplicates.

With denovo, there isn’t a baseline coordinate mapping to use, and gstacks isn’t the tool choice.

See the guide here

> [@Removing PCR duplicates in GStacks2](https://help.galaxyproject.org/t/removing-pcr-duplicates-in-gstacks2/16224/2):
>
> [Stacks: Stacks Manual](https://catchenlab.life.illinois.edu/stacks/manual/)

then, the FAQs link has [Stacks: Stacks: Frequently Asked Questions](https://catchenlab.life.illinois.edu/stacks/faq.php#format)

> ### What are the input and output data formats for Stacks?
> 
> In the _de novo_ case, data is read by the ustacks program and it currently can read either FASTA, FASTQ, or BAM formats. When a reference genome is available, aligned data is read by the gstacks program and either [SAM](http://www.htslib.org/doc/samtools.html) or BAM formats can be input.

The tool is in Galaxy as one of these

- **Stacks2: ustacks** Identify unique stacks

- **Stacks2: de novo map** the Stacks pipeline without a reference genome

- **Stacks: ustacks align** short reads into stacks

- **Stacks: de novo map** the Stacks pipeline without a reference genome (denovo\_map.pl)(Galaxy Version 1.46.0)

The protocol is in a publication (paywalled 🙃) but we don’t have a dedicated Galaxy tutorial I can point you to. Instead, try searching online to see if anyone has broken this out if you can’t see the paper. The core steps will be about the same in Galaxy – the difference is usually just how to set the metadata such as datatypes, and these are all fastq, BAM, tabular datatypes, which are common across many tools.

If you are new to Galaxy, consider running though a Learning Pathway like this to get familiar with how to organize data and navigate around the interface. And, if you are already familiar with Bioinformatics analysis, you can simply consider this a reference. → [Learning Pathway: Introduction to Galaxy and Sequence analysis](https://training.galaxyproject.org/training-material/learning-pathways/intro-to-galaxy-and-genomics.html).

Hope this helps again! 🙂

---

_[View the full topic](https://help.galaxyproject.org/t/removing-pcr-duplicates-in-gstacks2-galaxy-eu/16393)._
