How to run VGP workflow with large files?

I’m trying to run VGP workflow for my genome assembly, but my HiFi data is large which is 230GB. It is about 60GB in compressed (.gz) format, but VGP workflow only accept uncompressed files. My total storage size is only 600GB, so after cutadapt step, the total size of the trim file and the original file can be more than 400GB. And when I create collection, the file appeared twice, so now the total file size is about 800GB. Is it possible to run VGP workflow for my data? How to reduce the storage usage when create collection?

Welcome @jinglin

I’m not sure where you are working, but the UseGalaxy servers all have extended quota options and you can also attach your own storage to avoid the upload/download steps.

That said, even if you have enough data storage space you might run into processing limitations with this much data! It is difficult to guess since the characteristics of the assembly and your parameter choices will influence this performance when using the public clusters.

The UseGalaxy.eu server could be the best public option since they can scale the cluster resources on demand and have 2 TB of extra scratch data storage space. If you are not working there yet. or if you run into memory or runtime limits at a different UseGalaxy server, you can try there next before looking into on demand cloud Galaxy server options that can scale based on commercial cluster resources.

I hope this helps but please let us know if it actually does or if you have followup questions! :slight_smile:

Thank you! I emailed Galaxy Australia and they extended my storage space.