# goseq error in length and/or count files?

**URL:** <https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878>\
**Category:** usegalaxy.eu support\
**Tags:** troubleshooting\
**Created:** [February 27, 2025, 1:28pm UTC](https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878 "2025-02-27T13:28:45Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![mars142](https://avatars.discourse-cdn.com/v4/letter/m/45deac/32.png) [@mars142](https://help.galaxyproject.org/u/mars142)\
**Post date:** [February 27, 2025, 1:28pm UTC](https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878/1 "2025-02-27T13:28:45Z")

</div>

Hello there,  
I did some RNAseq analysis, which actually worked well.  
Now I wanted to move on with DEG analysis. I took the result files from the former analysis and everything worked fine until the goseq analysis:

Error in `[.default`(summary(map), , 1) : incorrect number of dimensions

This happens either I manipulate my data or use the originals from the former feature count analysis.

I know that two coworkers also used these exact workflows without any issues.  
May you help with some expertise? Thanks 🙂

History:

> **[Galaxy](https://usegalaxy.eu/published/history?id=8feadcb5a63f33f6)**
>
> Galaxy is a community-driven web-based analysis platform for life science research.

Workflow invocation: [9dbd24ecdd3fcec8](https://usegalaxy.eu/workflows/invocations/9dbd24ecdd3fcec8)

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [February 27, 2025, 7:13pm UTC](https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878/2 "2025-02-27T19:13:35Z")

</div>

Welcome @mars142

Thanks for sharing your history, super helpful!

The problem is with the **Ensembl** gene identifiers – these have an extra version `.N` and the **goseq** tool doesn’t understand how to parse those correctly.

If you scroll down on the tool form, you’ll see the expected format, and those happen to use **Ensembl** too.

I’m guessing that your coworkers were either using a different annotation source, or they removed the versioning.

This recent topic here has more details (a different Bioconductor tool, but all are in R and work the same at a technical level). →

> [@Missing gene id labeling in limma](https://help.galaxyproject.org/t/missing-gene-id-labeling-in-limma/14772/13):
>
> Some of the Bioconductor tools do not understand the “dot” in identifiers. Sort of a gotcha but that is how the tools work everywhere. In short, R is interpreting the dot when it shouldn’t be. I think you can quote it but I haven’t tested that with every tool in this pipeline.
> 
> We have a few topics about it if you are curious. The solution can be to remove the `.N` part of the identifier. This is one example.
> 
> - [how to replace these ID with official gene names?](https://help.galaxyproject.org/t/how-to-replace-these-id-with-official-gene-names/726)
> 
> The FAQ here has a troubleshooting warning but it is easy to miss!
> 
> - [FAQ: Extended Help for Differential Expression Analysis Tools](https://training.galaxyproject.org/training-material/faqs/galaxy/analysis_differential_expression_help.html)
> - Sometimes these tools do not understand `transcript_id.N` and `gene_id.N` notation (where N is a version number).
> - This notation could be in fasta or tabular inputs.
> - Try [removing `.N` from all inputs](https://training.galaxyproject.org/training-material/search2.html?query=olympics), and check for the accidential creation of new duplicates!
> 
> This trips up people using the tools directly, too!
> 
> - [https://support.bioconductor.org/post/search/?query=version+numbers+on+identifiers](https://support.bioconductor.org/post/search/?query=version+numbers+on+identifiers)

* * *

**What to do**

- Add a step into the workflow that strips of the `.N` content in the tabular files before sending the data into **goseq**.
- Find an annotation file that already has the version stripped.
- More about genome data sources. → [Reference genomes at public Galaxy servers: GRCh38/hg38 example](https://help.galaxyproject.org/t/reference-genomes-at-public-galaxy-servers-grch38-hg38-example/11616)

Hope this helps! 🙂

---

<div class="post-metadata">

**Author:** ![mars142](https://avatars.discourse-cdn.com/v4/letter/m/45deac/32.png) [@mars142](https://help.galaxyproject.org/u/mars142)\
**Post date:** [March 5, 2025, 3:05pm UTC](https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878/3 "2025-03-05T15:05:50Z")

</div>

Hey, thanks for your help. Seemed to work. Now I just got the problem of having absolute zero differences in expression values, which should not be the case. I can not find the case for that. Maybe I switched data or parameters. But I did not find any differences, while going through the training material.  
For your information, I want to analyse 5 treated cell cultures against one control. So looking over all experiments, that were done before, there must be a difference in gene expression.

> **[Galaxy](https://usegalaxy.eu/published/history?id=8feadcb5a63f33f6)**
>
> Galaxy is a community-driven web-based analysis platform for life science research.

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [March 5, 2025, 10:06pm UTC](https://help.galaxyproject.org/t/goseq-error-in-length-and-or-count-files/14878/4 "2025-03-05T22:06:58Z")

</div>

Hi @mars142

It is hard to see what is going on with this result. But I do see that you had problems linking back in the annotation. You could switch to using the UCSC version of this annotation instead. It already has the simplified Ensembl genes/transcripts and will work with all these tools without extra manipulations.

- [https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/genes/](https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/genes/)

Other than that, maybe only having one control sample is the root issue.

There is probably a lot of discussion about this online, but we have one captured in our FAQ here as a starting place.

- [FAQ: Extended Help for Differential Expression Analysis Tools](https://training.galaxyproject.org/training-material/faqs/galaxy/analysis_differential_expression_help.html)
- See → Differential expression tools all require sample count replicates. [Rationale from two of the DEseq tool authors](https://www.seqanswers.com/forum/bioinformatics/bioinformatics-aa/26388-deseq2-without-biol-replicates).

Hope this helps! 🧑‍🔬
