# filtering a single-end fastq.gz collection by number of reads

**URL:** <https://help.galaxyproject.org/t/filtering-a-single-end-fastq-gz-collection-by-number-of-reads/10777>\
**Category:** Uncategorized\
**Tags:** text-manipulation, collections\
**Created:** [September 12, 2023, 1:54pm UTC](https://help.galaxyproject.org/t/filtering-a-single-end-fastq-gz-collection-by-number-of-reads/10777 "2023-09-12T13:54:29Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![microfuge](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/microfuge/32/3587_2.png) [@microfuge](https://help.galaxyproject.org/u/microfuge)\
**Post date:** [September 12, 2023, 1:54pm UTC](https://help.galaxyproject.org/t/filtering-a-single-end-fastq-gz-collection-by-number-of-reads/10777/1 "2023-09-12T13:54:29Z")

</div>

Hi ,

We have a single end fastq collection with thousands of files and want to keep only files having more than 200 reads in it. Is there a way to do this with existing tools in Galaxy?

Thanks!  
Saurabh

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [September 12, 2023, 7:49pm UTC](https://help.galaxyproject.org/t/filtering-a-single-end-fastq-gz-collection-by-number-of-reads/10777/2 "2023-09-12T19:49:35Z")

</div>

Hi @microfuge

There isn’t an exact tool but you could string together multiple tools into a workflow to do this.

Process will be something like: count up the number of lines per file, filter on the line count values, capture the identifiers from elements/files that pass the filter, then filter the original collection with those.

The tools you will need will be covered in these:

- [Data Manipulation Olympics](https://training.galaxyproject.org/training-material/topics/introduction/tutorials/data-manipulation-olympics/tutorial.html)
- [Dataset Collections](https://training.galaxyproject.org/training-material/search2?query=collection)

---

<div class="post-metadata">

**Author:** ![microfuge](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/microfuge/32/3587_2.png) [@microfuge](https://help.galaxyproject.org/u/microfuge)\
**Post date:** [September 13, 2023, 9:57am UTC](https://help.galaxyproject.org/t/filtering-a-single-end-fastq-gz-collection-by-number-of-reads/10777/3 "2023-09-13T09:57:18Z")

</div>

Thanks @jennaj

I used the tools toolshed.g2.bx.psu.edu/repos/iuc/seqkit\_stats/seqkit\_stats/2.2.0+galaxy0 and then “Collapse collection into a single dataset” to obtain a tsv file, kept the required entries with awk and then “filter collection” .
