# Using CHM13v2.0 T2T and mm39 gene annotations for RNASeq Analysis

**URL:** <https://help.galaxyproject.org/t/using-chm13v2-0-t2t-and-mm39-gene-annotations-for-rnaseq-analysis/16477>\
**Category:** usegalaxy.eu support\
**Tags:** transcriptomics, reference-annotation, reference-genome, tool-help, deseq2, featurecounts\
**Created:** [November 6, 2025, 12:31pm UTC](https://help.galaxyproject.org/t/using-chm13v2-0-t2t-and-mm39-gene-annotations-for-rnaseq-analysis/16477 "2025-11-06T12:31:15Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [November 6, 2025, 9:01pm UTC](https://help.galaxyproject.org/t/using-chm13v2-0-t2t-and-mm39-gene-annotations-for-rnaseq-analysis/16477/2 "2025-11-06T21:01:33Z")

</div>

Welcome @jankat

For RNA-seq, using the human hg38 genome will be the “most current” genome assembly. The annotation for the T2T assembly for genes/transcripts _ **is** _ the hg38 annotation. The data is _mapped over_ to the T2T coordinates. You will not need that and the extra exposed regions are unlikely to help and can harm.

For the mouse, yes, mm39 is the more current version, and annotation is available for it.

### How to use these

The annotation’s features are described by coordinates on the chromosome bases. Different assemblies have different bases! So try to avoid mixing up files between assembly versions. You will also want to use the same annotation throughout an analysis for all steps since the features are also labeled in specific ways.

If you later want to use a different assembly, any step you already did that uses that assembly’s bases or coordinate system will need to be rerun.

Same for annotation. If you want to use a different GTF, all steps that use any annotation source (whether supplied by you, or built into a tool), will need to be rerun using it.

### Guides

- Reference data assembly choices. → [Reference genomes at public Galaxy servers: GRCh38/hg38 example](https://help.galaxyproject.org/t/reference-genomes-at-public-galaxy-servers-grch38-hg38-example/11616)

- Getting the data organized for your target tools. → [FAQ: Extended Help for Differential Expression Analysis Tools](https://galaxyproject.github.io/training-material/faqs/galaxy/analysis_differential_expression_help.html)

- With much more at this forum under #transcriptomics and #reference-genome #reference-annotation – plus search with tools names such as #featurecounts

### Sources

The genomes for **hg38** and **mm10/mm39** were all originally hosted by UCSC (in coordination with NCBI), and Gencode uses the same chromosome/coordinate scheme, so you could get the annotation from either place and these reference annotation GTFs will work with the reference assembly genomes natively indexed in Galaxy.

The example here is for hg38, but mm39 will work the same.

> [@How to find the reference transcriptome for analysis tools: GRCh38/hg38 example](https://help.galaxyproject.org/t/how-to-find-the-reference-transcriptome-for-analysis-tools-grch38-hg38-example/12506/2):
>
> For the UCSC hg38 reference genome indexed in Galaxy, a reference annotation GTF and reference transcriptome fasta can be sources from at least these two places:
> 
> **Gencode**
> 
> - [GENCODE - Human Release 49](https://www.gencodegenes.org/human/)
> - get the first in the list of GTFs, and the first in the list of Fasta
> - double check the formatting. You might need to standardize the fasta with “NormalizeFasta” (I can’t remember if this is needed) and I would remove GTF headers too (some tools might have a problem with them). The FAQ above has instructions for these.
> 
> **UCSC**
> 
> These two are a match, and have standard human Gene Symbol and RefSeq transcript identifiers.
> 
> - [https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/genes/hg38.refGene.gtf.gz](https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/genes/hg38.refGene.gtf.gz)
> - [https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/refMrna.fa.gz](https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/refMrna.fa.gz)
> - The GTF will be ready to use after **Upload** , and the fasta will need to be uncompressed (under the pencil icon) then run through **NormalizeFasta** to strip out the extra characters.
> 
> Technically, any of the reference annotation GTFs in their Downloads area are based on “Gene and Gene Predictions” tracks also represented in the Table browser (or main Browser). This means you can extract a reference transcriptome fasta from the Table browser.

This means for mouse at UCSC, this is where you’ll be looking:

> [@RNA Seq, error alignment after trimming](https://help.galaxyproject.org/t/rna-seq-error-alignment-after-trimming/8791/5):
>
> **mm10** here [Index of /goldenPath/mm10/bigZips/genes](https://hgdownload.soe.ucsc.edu/goldenPath/mm10/bigZips/genes/)  
> **mm39** here [Index of /goldenPath/mm39/bigZips/genes](https://hgdownload.soe.ucsc.edu/goldenPath/mm39/bigZips/genes/)
> 
> The path is similar for other UCSC genome databases indexed for mapping tools. Copy the URL for the GTF file, paste that into the Upload tool, and leave all settings at default. The GTF will automatically uncompress and be ready to use without other formatting changes.
> 
> The **NCBIRefSeq GTF** will be an appropriate choice.

And Gencode for mouse, will be here

> [@RNA Star and mouse Ensembl GRCm39 problem](https://help.galaxyproject.org/t/rna-star-and-mouse-ensembl-grcm39-problem/9129/2):
>
> **GENCODE - Release M10 (GRCm38)** → [GENCODE - Mouse Release M10](https://www.gencodegenes.org/mouse/release_M10.html)
> 
> **GENCODE - Release M38 (GRCm39)** → [GENCODE - Mouse Release M38](https://www.gencodegenes.org/mouse/)
> 
> The **Basic gene annotation GTF** will be an appropriate choice.

* * *

* * *

Learning how to source and prepare your reference data is a very good skill to have. Your results will only be as good as the other data it is built upon!

Please give this a try, and if you get stuck, please ask and we can pull out the exact URLs for the data. We’ll need to know which source, which assembly, which annotation track, and which database key you decide to use for each.

Let’s start there! 🙂

---

_[View the full topic](https://help.galaxyproject.org/t/using-chm13v2-0-t2t-and-mm39-gene-annotations-for-rnaseq-analysis/16477)._
