Research questionHow can shared GPUs schedule concurrent heterogeneous AI inference without combinatorial profiling as workloads change?Concurrent models contend for shared GPU resources, making each joint configuration expensive to characterize. Runtime schedulers must still choose configurations quickly as workload conditions change.