Prefill and decode are two different workloads sharing one GPU. Their resource demands differ, and an unchunked long prefill can delay active decodes. Profile the actual model and batch shape before choosing a scheduling change.
an account raises the per-topic limit · no card