# Taxonomy workflows: reads to taxa

**URL:** <https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541>\
**Category:** Uncategorized\
**Tags:** metagenomics, iwc-workflows\
**Created:** [December 2, 2025, 1:20pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541 "2025-12-02T13:20:25Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Adam\_Hillier](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/adam_hillier/32/6564_2.png) [@Adam\_Hillier](https://help.galaxyproject.org/u/Adam_Hillier)\
**Post date:** [December 2, 2025, 1:20pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/1 "2025-12-02T13:20:25Z")

</div>

Hi Jemma,

I have some samples being sequenced. These are for soil fungi eDNA metarcoding.

On NCBI I have found and downloaded a FASTQ file (2MB) of the type I will be using.

Sequences will be about 300bp long.

I would like to use this to set up the following bioinformatics workflow so I am ready to go when I get my FASTQ files.

1. read in Illumina MiSeq FASTQ files (the files will already be demultiplexed)
2. trim sequences (if needed)
3. filter them e.g. reject low quality score) (if necessary).
4. Cluster in OTUs using appropriate parameters
5. Match OTUs using appropriate parameters to a suitable UK fungi sequence database
6. Output the taxa list to a spreadsheet

I would like to generate QC charts at each stage.

Might be some more steps but that would be a good start.

This is a pretty standard workflow and there must be a workflow like this already in Galaxy for me to copy/adapt.

How would I find it ?

How do I start building and testing a workflow ?

Adam

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [December 3, 2025, 7:56pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/2 "2025-12-03T19:56:18Z")

</div>

Hello @Adam_Hillier

Glad to learn you are proceeding with your project! 🧑‍🔬

There are a few primary places to source workflows for Galaxy.

### IWC – Production quality workflows

These are curated, so if you can find what you need here, that would be preferred! Each has been optimized for large batch stream processing. This catalog is newer and growing and has stricter community standards.

- **[Galaxy IWC - Workflow Library](https://iwc.galaxyproject.org/)**

### Galaxy Hub – Public Workflows

This is the “meta” search. You’ll find workflows from the GTN trainings, WorkflowHub, and the Public Workflows available from the communities at the UseGalaxy\* servers.

- **[GTN Pan-Galactic Workflow Search](https://training.galaxyproject.org/training-material/workflows/list)**

* * *

* * *

A workflow from either can be customized further, too. I would be pretty common to break out an analysis like yours into two or three distinct module workflows. Then scientists could run them separately or nest as _subworkflows_ into a single master workflow that does everything with a bit more customization ([reference data preparation](https://help.galaxyproject.org/t/faq-ncbi-reference-data/15720), [intermediate file offloading](https://help.galaxyproject.org/t/data-storage-choices-when-using-workflows/14372), [workflow reports](https://galaxyproject.github.io/training-material/topics/galaxy-interface/tutorials/workflow-reports/tutorial.html)).

I didn’t find an IWC workflow for eDNA specifically and one of the training quality workflows from the GTN is probably too simple for your needs (no clustering). The other, using **Obitools** , will work best at one of the “Available at these Galaxies” servers for now – [UseGalaxy.eu](http://UseGalaxy.eu) would be a good choice. 🙂

- [Hands-on: Metabarcoding/eDNA through Obitools / Metabarcoding/eDNA through Obitools / Ecology](https://galaxyproject.github.io/training-material/topics/ecology/tutorials/Obitools-metabarcoding/tutorial.html)
- Note that the training version of workflows may have disconnected steps (on purpose!) and this one does. But, you’ll probably want to be starting at step 4 anyway since you already have your fastq reads prepared. I would suggest deleting the training initial steps, adding in your [collection input](https://help.galaxyproject.org/t/why-collections/16411), then connect.

Hope this helps to get things started! 🙂

---

<div class="post-metadata">

**Author:** ![Adam\_Hillier](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/adam_hillier/32/6564_2.png) [@Adam\_Hillier](https://help.galaxyproject.org/u/Adam_Hillier)\
**Post date:** [December 22, 2025, 6:01pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/3 "2025-12-22T18:01:38Z")

</div>

Hi,

I have got the Wolf Diet eDNA metadata set and workflow and have run rhat.

It almost worked. But it failed at one step and I am struggling to diagnose. Pretty sure if I get that fixed it will then run to completion.

If I can get that to work I imagine I will be able to adapt the workflow for my own data.

Adam

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [December 23, 2025, 10:02pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/4 "2025-12-23T22:02:11Z")

</div>

Great! Glad you learn you have made progress! 🙂

If you think the issue was technical, and would like some feedback, you are welcome to post back a share link to the history (or better, the workflow invocation – see the Share button in the top menu bar of that view). Maybe we can help to solve it here?

If you would rather share in a private thread, we can do that, but it will be harder to get feedback from our developers (if needed). You can decide.

---

<div class="post-metadata">

**Author:** ![Adam\_Hillier](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/adam_hillier/32/6564_2.png) [@Adam\_Hillier](https://help.galaxyproject.org/u/Adam_Hillier)\
**Post date:** [December 29, 2025, 9:21pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/5 "2025-12-29T21:21:27Z")

</div>

HI again,

I have managed to read in one of the WolfDiet fastq files and checked that it’s ok.

I have blasted the fastq file against the reference database.

So I have a list of sequences which I presume are all matched to a species name.

What I am struggling with now is how to add family, order, class and phylum to each match sequence.

Once I can do that I am pretty much there.

Adam

---

<div class="post-metadata">

**Author:** ![Adam\_Hillier](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/adam_hillier/32/6564_2.png) [@Adam\_Hillier](https://help.galaxyproject.org/u/Adam_Hillier)\
**Post date:** [December 29, 2025, 11:51pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/6 "2025-12-29T23:51:56Z")

</div>

Hi,

I think I am missing some tools such as Taxonkit and NCBI Taxomomy.

Any ideas why I can’t see them in tools ?

These will allow me to use a column eith taxonomic ID to produce other columns of taxonomic lineage (phylum, class, order, family, genus, species.

Adam

---

<div class="post-metadata">

**Author:** ![Adam\_Hillier](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/adam_hillier/32/6564_2.png) [@Adam\_Hillier](https://help.galaxyproject.org/u/Adam_Hillier)\
**Post date:** [December 30, 2025, 5:14pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/7 "2025-12-30T17:14:32Z")

</div>

HI team,

Galaxy needs a published metabarcode database creation and BLAST workflow that works. I am happy to do that.

The Wolf Diet workflow writeup talks about **ecoPCR** and **ecoTag** which I cant find in Galaxy tools.

I have also read about **NCBI Taxomony** and **Taxonkit** which should be on Galaxy but I can’t find these either. Do I have access to only some tools ?

Without these I can only BLAST my depelicated metabarcode fasta sequences against the full NCBI nucleotide database (see below) which will take much much longer.

Adam

 ![3C64A6A5-5726-4C04-A5BF-20B7C2E62E4F.jpeg](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/2X/3/3d2ec0e0667e72ed4284d6253914df5086e4355d.jpeg)

’

 ![99390E01-24CF-48F1-8143-7FE3068D4818_4_5005_c.jpeg](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/2X/1/1bf54e0c1cf8785978edb7ca712ebd0435b14215.jpeg)

---

<div class="post-metadata">

**Author:** ![Adam\_Hillier](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/adam_hillier/32/6564_2.png) [@Adam\_Hillier](https://help.galaxyproject.org/u/Adam_Hillier)\
**Post date:** [December 31, 2025, 5:40pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/8 "2025-12-31T17:40:23Z")

</div>

Hi all,

Update on my Wolf Diet workflow trial.

I got the BLAST to run reasonably quickly using the full NCBI NT (15 Aug 2024) database. Not ideal perhaps but it worked.

I then merged in the taxonomy lineage by downloading a file with TAXIDs (output from BLAST) and running taxonKit in my terminal and then uploading the file back into Galaxy. ChatGPT help me do this.

This allowed me to produce a Krona Chart

 ![53E116C6-BA6F-40B3-90F5-3D1BACA3C0A5.jpeg](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/2X/7/7d1845c19186a9324b8edc7a8c3ecd11eab611fd.jpeg)

I then cut and paste the species list into ChatGPT and asked it to tell me where the samples were from (see below)

Adam

## 1️⃣ **China** — **strongest match**

China fits **both wolves and many realistic prey species** in your data.

**Wolf prey from your list found in China**

- **Capreolus pygargus** (Siberian roe deer)

- **Cervus elaphus** (red deer)

- **Cervus nippon** (sika deer)

- **Cervus albirostris** (white-lipped deer)

- **Elaphodus cephalophus** (tufted deer)

- **Procapra gutturosa / picticaudata / przewalskii** (gazelles)

- **Pantholops hodgsonii** (Tibetan antelope)

- **Marmota sibirica / himalayana** (marmots)

**Why this works**

- Wolves are native to **northern & western China**

- These are **documented wolf prey** in steppe, plateau, and forest systems

- No need to invoke introductions or zoos

---

<div class="post-metadata">

**Author:** ![Adam\_Hillier](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/adam_hillier/32/6564_2.png) [@Adam\_Hillier](https://help.galaxyproject.org/u/Adam_Hillier)\
**Post date:** [December 31, 2025, 5:45pm UTC](https://help.galaxyproject.org/t/taxonomy-workflows-reads-to-taxa/16541/9 "2025-12-31T17:45:14Z")

</div>

Hi all,

I now feel ready to use Galaxy to process and analyse my own data.

It would be good if **Taxonkit** was in Galaxy. Is that possible ? Then the whole workflow could be done in Galaxy.

Is there a published paper about the wolf diets so I can read it and compare with my findings ?

Thanks

Adam
