# How to merge or generate FASTA file with the same chromosome

**URL:** <https://help.galaxyproject.org/t/how-to-merge-or-generate-fasta-file-with-the-same-chromosome/190>\
**Category:** usegalaxy.org support\
**Tags:** devops-administration, quality-control\
**Created:** [December 12, 2018, 5:21pm UTC](https://help.galaxyproject.org/t/how-to-merge-or-generate-fasta-file-with-the-same-chromosome/190 "2018-12-12T17:21:44Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nattawat\_Chaiyawong](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/nattawat_chaiyawong/32/385_2.png) [@Nattawat\_Chaiyawong](https://help.galaxyproject.org/u/Nattawat_Chaiyawong)\
**Post date:** [December 12, 2018, 5:21pm UTC](https://help.galaxyproject.org/t/how-to-merge-or-generate-fasta-file-with-the-same-chromosome/190/1 "2018-12-12T17:21:44Z")

</div>

Hi

I try to generate FASTA file from BAM file. I followed the protocol that I got from [https://usegalaxy.org/u/antunderwood/w/bam-to-fasta](https://usegalaxy.org/u/antunderwood/w/bam-to-fasta) However, after getting FASTA file, the file does not merge the same chromosome. The file shows the same chromosome with several fragments.

 ![27](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/1X/0b257027d605825d0d962b824e169e7876ef5f9d.png) For example Chr1 0-100, Chr1 101-200, Chr1 201-300…Chr 1001-2000… I need to merge all dataset with the same chromosome together which will make me easy to analyze the data. Such as, Chr1 0-10XXXX, Chr2 0-10XXXX, Chr3 0-10XXXX… Could you give me some suggestions about this problem? Thank you so much.

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [December 12, 2018, 9:08pm UTC](https://help.galaxyproject.org/t/how-to-merge-or-generate-fasta-file-with-the-same-chromosome/190/2 "2018-12-12T21:08:40Z")

</div>

Hello,

Examine the fasta identifiers in your output and note that they have regions included. Something like:

`>chr1:0-100`  
`ATCGATCGATCGATCGATCG`

When what you want seems to be:

`>chr1 0-100`  
`ATCGATCGATCGATCGATCG`

Use one of the **Replace Text** tools to modify the “\>” lines. Just be aware that renaming the sequences without the coordinates included in the fasta identifiers will create duplicates.

FAQ for Fasta format: [https://galaxyproject.org/learn/datatypes/#fasta](https://galaxyproject.org/learn/datatypes/#fasta)

---

<div class="post-metadata">

**Author:** ![Nattawat\_Chaiyawong](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/nattawat_chaiyawong/32/385_2.png) [@Nattawat\_Chaiyawong](https://help.galaxyproject.org/u/Nattawat_Chaiyawong)\
**Post date:** [December 12, 2018, 9:44pm UTC](https://help.galaxyproject.org/t/how-to-merge-or-generate-fasta-file-with-the-same-chromosome/190/3 "2018-12-12T21:44:15Z")

</div>

> [@jennaj](#):
>

Hello,

Thank you for your comment. But, I still don’t understand the answer you provided. According to my question, my FASTA file shows Chromosome like Py17x\_01\_v3:0-1073, Py17x\_01\_v3:1102-1255, Py17x\_01\_v3:1259-1587 … Py17x\_01\_v3:813740-815147 (figure above). But I need the FASTA file show Py17x\_01\_v3:0-815147 for the Chr1 and Py17x\_02\_v3:0-xxxxxxxx for Chr2 … until Chr14. Because I want to see all sequence in the same window when I analyze the data using IGV. I don’t want to separate the data with the same chromosome. I am sorry if my question is not clear. Could you explain how I can modify it in detail, please? Thank you so much.

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [December 12, 2018, 10:02pm UTC](https://help.galaxyproject.org/t/how-to-merge-or-generate-fasta-file-with-the-same-chromosome/190/4 "2018-12-12T22:02:29Z")

</div>

Look at your custom genome/BAM hit data. My guess is that it has chromosomes named like “Py17x\_01\_v3”, not “chr1”.

You’ll need to start off by mapping against a genome that has chromosomes named in a way that you want to use or group by in later steps, otherwise, the identifiers/coordinates won’t be a match.

---

<div class="post-metadata">

**Author:** ![Nattawat\_Chaiyawong](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/nattawat_chaiyawong/32/385_2.png) [@Nattawat\_Chaiyawong](https://help.galaxyproject.org/u/Nattawat_Chaiyawong)\
**Post date:** [December 12, 2018, 10:24pm UTC](https://help.galaxyproject.org/t/how-to-merge-or-generate-fasta-file-with-the-same-chromosome/190/5 "2018-12-12T22:24:58Z")

</div>

Yes, the chromosome name is Py17x\_01\_v3, Py17x\_02\_v3…, Py17x\_14\_v3 (14 chromosomes). For my understanding, I need to start with mapping against a genome (same chromosome name) to generate the BAM file and then convert it to FASTA file, right? The reason that I want to generate the FASTA file is I will use this FASTA file for the template (reference genome) for my further analysis (mapping this reference genome with my unknown sample). Thank you.

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [December 12, 2018, 11:34pm UTC](https://help.galaxyproject.org/t/how-to-merge-or-generate-fasta-file-with-the-same-chromosome/190/6 "2018-12-12T23:34:26Z")

</div>

I’m not quite sure I understand but it sounds like you are trying to build a new reference transcriptome (or other “-ome”). If so, it may help to review the Galaxy tutorials.

> [@Troubleshooting resources for errors or unexpected results](https://help.galaxyproject.org/t/troubleshooting-resources-for-errors-or-unexpected-results/42/1):
>
> [Galaxy Training Network Tutorials](https://galaxyproject.github.io/training-material/): Some [GTN](https://galaxyproject.org/teach/gtn/) tutorials are appropriate for [Galaxy Main](https://usegalaxy.org) and some are not. Where you can run each is noted per tutorial – click on the _Galaxy instances_ gear icon ![galaxyserversGTN](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/1X/1287f50cfcdaa1aad750fe323f96995260ac4947.png) to review the [Public Galaxy](https://galaxyproject.org/use/) server choices. If a tutorial is supported by a pre-configured [Galaxy Docker](https://hub.docker.com/r/bgruening/galaxy-stable/) training image, instructions for how to get it will be listed below the tutorial listings, per category.

_Note_: It is technically possible to rename identifiers in any dataset but some are more difficult to transform than others. And all datasets used together in visualization or other analysis steps need to have the identifiers (that represent the exact same underlying data) modified with a precision method that fits the different datatypes involved. I wouldn’t recommend that anyone attempts modifications like this unless they already know how to do it and are able to detect plus troubleshoot any problems that might come up. The steps are too complicated (especially with BAM data) and much can go wrong. This is why I think that starting over with the identifiers you want to use at the beginning, then working through your analysis with consistent data for all steps, is the best approach. All will go much smoother!

Since you are starting over, this might be a good time to consider updating the workflow you shared. It uses older versions of a few tools. Using the latest versions of tools is always better. Tools can be updated within the workflow editor.
