# BAM file from HISAT2 to fastq for Genome assembly

**URL:** <https://help.galaxyproject.org/t/bam-file-from-hisat2-to-fastq-for-genome-assembly/15031>\
**Category:** usegalaxy.eu support\
**Tags:** bedtools, fastq-format\
**Created:** [March 17, 2025, 3:05pm UTC](https://help.galaxyproject.org/t/bam-file-from-hisat2-to-fastq-for-genome-assembly/15031 "2025-03-17T15:05:46Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![John\_Kim](https://avatars.discourse-cdn.com/v4/letter/j/4491bb/32.png) [@John\_Kim](https://help.galaxyproject.org/u/John_Kim)\
**Post date:** [March 17, 2025, 3:05pm UTC](https://help.galaxyproject.org/t/bam-file-from-hisat2-to-fastq-for-genome-assembly/15031/1 "2025-03-17T15:05:46Z")

</div>

Hi friends!,

I would like to ask if you know of any tools to convert BAM alignment files (which I generated in Galaxy using HISAT2) into FASTQ files, suitable for genome assembly. I used `bedtools bamtofastq`, but I received a warning. Is this normal?

 ![image](https://us1.discourse-cdn.com/flex020/uploads/galaxy/original/2X/b/bbeb9a3e79d86104e64dc90df55e1187c04da1d0.png)

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [March 17, 2025, 5:06pm UTC](https://help.galaxyproject.org/t/bam-file-from-hisat2-to-fastq-for-genome-assembly/15031/2 "2025-03-17T17:06:39Z")

</div>

Hi @John_Kim

The warning we can see is about unpaired reads inside the original BAM. Most assembly tools that accept paired reads will expect those to be intact pairs. Your fastq files will probably need to be filtered now.

However, you were mapping to filter the reads, correct? If you filtered the BAM first, then extracted the reads, you might prefer that set of reads for assembly. Right now any sequence with any hit is being output.

Search the tool panel with “filter bam” to find the tool choices. Proper pairs, primary alignments, removing unmapped, and some minimum mapQ values are common choices for many analysis paths but which to use are your choices.

To see the full logs and warning, review the `stderr/stdout` messages from the tool on the Job Details view ([using the i-info icon](https://training.galaxyproject.org/training-material/faqs/galaxy/datasets_icons.html)).

Hope this helps and we can follow up more! 🙂

---

<div class="post-metadata">

**Author:** ![John\_Kim](https://avatars.discourse-cdn.com/v4/letter/j/4491bb/32.png) [@John\_Kim](https://help.galaxyproject.org/u/John_Kim)\
**Post date:** [March 18, 2025, 3:46am UTC](https://help.galaxyproject.org/t/bam-file-from-hisat2-to-fastq-for-genome-assembly/15031/3 "2025-03-18T03:46:22Z")

</div>

Good day @jennaj!

I hope you’re doing well. I’m going to retry the steps, starting with the BAM file generated by HISAT2. Next, I plan to filter the BAM file using `Filter BAM`, correct? Considering my objective is solely to assemble reads and extract the `.fastq` files from the aligned BAM, what parameters do you recommend I use?

Regards,  
John\_Kim

---

<div class="post-metadata">

**Author:** ![jennaj](https://sea2.discourse-cdn.com/flex020/user_avatar/help.galaxyproject.org/jennaj/32/27_2.png) [@jennaj](https://help.galaxyproject.org/u/jennaj)\
**Post date:** [March 18, 2025, 6:40pm UTC](https://help.galaxyproject.org/t/bam-file-from-hisat2-to-fastq-for-genome-assembly/15031/4 "2025-03-18T18:40:56Z")

</div>

Hi @John_Kim

Some of this depends on why you were filtering originally.

You could try different parameters sets, then compare the assemblies, and decide that way? At a minimum you want intact pairs since that is what the assembly tools will expect, and “proper pairs” will ensure those pairs are mapping to the same place, and not other random places in the genome.

Low quality reads – both low technical quality (actual quality scores) and low scientific quality (however you define that) – will increase the chances of complications with any assembly method.

We have assembly tutorials here if interested:

> **[Assembly / Tutorial List](https://training.galaxyproject.org/training-material/topics/assembly/)**
>
> DNA sequence data has become an indispensable tool for Molecular Biology & Evolutionary Biology. Study in these fields now require a genome sequence to work from. We call this a 'Reference Sequence.' We need to build a reference for each species....

> **[Genome Annotation / Tutorial List](https://training.galaxyproject.org/training-material/topics/genome-annotation/)**
>
> Genome annotation is a multi-level process that includes prediction of protein-coding genes, as well as other functional genome units such as structural RNAs, tRNAs, small RNAs, pseudogenes, control regions, direct and inverted repeats, insertion...

You can also find these same tutorials linked on the bottom of tool forms. Reviewing those will filter by that tool.

Hope this helps! 🙂
