I investigated a bit more and could bring it down to an issue of the resubmission and the global limits.
execution:
default: medium_memory_and_high_cpu
environments:
medium_memory_and_high_cpu:
runner: slurm
native_specification: '--cpus-per-task=10 --mem=400'
resubmit:
- condition: memory_limit_reached
environment: high_memory_and_cpu
high_memory_and_cpu:
runner: slurm
native_specification: '--cpus-per-task=10 --mem=800'
resubmit:
- condition: memory_limit_reached
environment: ultra_high_memory
ultra_high_memory:
runner: slurm
native_specification: '--cpus-per-task=10 --mem=300000'
limits:
- type: destination_user_concurrent_jobs
id: medium_memory_and_high_cpu
value: 5
- type: destination_user_concurrent_jobs
id: high_memory_and_cpu
value: 3
- type: destination_user_concurrent_jobs
id: ultra_high_memory
value: 1
In this setting, if the number jobs as given per global limit e.g. for ultra_high_memory already 1 job, is submitted, it ‘somehow’ blocked from resubmission, and we end up in the endless resubmission loop. If I increase the limit to e.g. 2, than one job can be rescheduled, but we run into the same issue if we have 2 jobs which need to be resubmitted.
chatGPT proposed to move the limits into the environment definition, but that also does not work. It further wanted to add handlers, but also no success.
Anyone knowns here something? I am very sure, that this setup worked absolutly fine in Galaxy version 24.2, and only since the upgrade to 25.0 I see this issue.