# goseq error in length and/or count files?

**URL:** <https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878>\
**Category:** usegalaxy.eu support\
**Tags:** troubleshooting\
**Created:** [February 27, 2025, 1:28pm UTC](https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878 "2025-02-27T13:28:45Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [February 27, 2025, 7:13pm UTC](https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878/2 "2025-02-27T19:13:35Z")

</div>

Welcome @mars142

Thanks for sharing your history, super helpful!

The problem is with the **Ensembl** gene identifiers – these have an extra version `.N` and the **goseq** tool doesn’t understand how to parse those correctly.

If you scroll down on the tool form, you’ll see the expected format, and those happen to use **Ensembl** too.

I’m guessing that your coworkers were either using a different annotation source, or they removed the versioning.

This recent topic here has more details (a different Bioconductor tool, but all are in R and work the same at a technical level). →

> [@Missing gene id labeling in limma](https://help.galaxyproject.org/t/missing-gene-id-labeling-in-limma/14772/13):
>
> Some of the Bioconductor tools do not understand the “dot” in identifiers. Sort of a gotcha but that is how the tools work everywhere. In short, R is interpreting the dot when it shouldn’t be. I think you can quote it but I haven’t tested that with every tool in this pipeline.
> 
> We have a few topics about it if you are curious. The solution can be to remove the `.N` part of the identifier. This is one example.
> 
> - [how to replace these ID with official gene names?](https://help.galaxyproject.org/t/how-to-replace-these-id-with-official-gene-names/726)
> 
> The FAQ here has a troubleshooting warning but it is easy to miss!
> 
> - [FAQ: Extended Help for Differential Expression Analysis Tools](https://training.galaxyproject.org/training-material/faqs/galaxy/analysis_differential_expression_help.html)
> - Sometimes these tools do not understand `transcript_id.N` and `gene_id.N` notation (where N is a version number).
> - This notation could be in fasta or tabular inputs.
> - Try [removing `.N` from all inputs](https://training.galaxyproject.org/training-material/search2.html?query=olympics), and check for the accidential creation of new duplicates!
> 
> This trips up people using the tools directly, too!
> 
> - [https://support.bioconductor.org/post/search/?query=version+numbers+on+identifiers](https://support.bioconductor.org/post/search/?query=version+numbers+on+identifiers)

* * *

**What to do**

- Add a step into the workflow that strips of the `.N` content in the tabular files before sending the data into **goseq**.
- Find an annotation file that already has the version stripped.
- More about genome data sources. → [Reference genomes at public Galaxy servers: GRCh38/hg38 example](https://help.galaxyproject.org/t/reference-genomes-at-public-galaxy-servers-grch38-hg38-example/11616)

Hope this helps! 🙂

---

_[View the full topic](https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878)._
