# How to remove duplicates in a concatenated paired dataset?

**URL:** <https://help.galaxyproject.org/t/how-to-remove-duplicates-in-a-concatenated-paired-dataset/6687>\
**Category:** usegalaxy.org.au support\
**Tags:** workflow, metagenomics, mothur\
**Created:** [September 16, 2021, 6:54am UTC](https://help.galaxyproject.org/t/how-to-remove-duplicates-in-a-concatenated-paired-dataset/6687 "2021-09-16T06:54:19Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![annalohning](https://avatars.discourse-cdn.com/v4/letter/a/ecb155/32.png) [@annalohning](https://help.galaxyproject.org/u/annalohning)\
**Post date:** [September 16, 2021, 6:54am UTC](https://help.galaxyproject.org/t/how-to-remove-duplicates-in-a-concatenated-paired-dataset/6687/1 "2021-09-16T06:54:19Z")

</div>

Hi,  
Am new to Galaxy but have processed my paired end sequences to date with mothur. Would love to try Galaxy. I’ve uploaded the fastq files my local dir and made pairs (85) into a dataset. However on inspection some did not transfer (should have 89 pairs) and one has an odd filename.

My queries is, can I  
(a) delete a pair in a dataset and/or  
(b) add additional pairs to an existing dataset?

I’ve tried uploading the missing .fastq files, created another dataset of pairs (6) and [concatentated](https://usegalaxy.org/datasets/bbd44e69cb8906b5edd35e5df082492b/display?to_ext=fastqsanger) them but there are duplicates that must be removed.

Given the above, is it better to simply start again (even though it took days to upload all the files?)

Many thanks for the assistance  
Anna
