Both workloads are legitimate and they want opposite things from the same GPU. Combine tenant rate limits with fair admission and measured capacity reservations; separate fleets are an option when shared scheduling cannot meet the targets.
an account raises the per-topic limit · no card