Get Started
Research questionHow can shared GPUs schedule concurrent heterogeneous AI inference without combinatorial profiling as workloads change?Concurrent models contend for shared GPU resources, making each joint configuration expensive to characterize. Runtime schedulers must still choose configurations quickly as workload conditions change.
AI
Inference Optimization
Technology
Latest papersRecent research connected to this question, newest first.MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUsThe evidence covers concurrent LLM and vision workloads with 2–6 co-running models, including eight dynamic workload scenarios and scheduling configurations on a shared GPU. Broader hardware, model, and workload coverage is not specified.research paper · Sep 2, 2026
Related questions
How can multi-GPU serving meet chunk-latency targets for stateful, bursty video-generation sessions?How can complex neural-network graphs be mapped across heterogeneous SoCs to balance inference latency and throughput?How can idle inference resources reduce scarce-GPU training cost without biasing gradient estimates?How can enterprises consolidate heterogeneous internal workloads onto one self-hosted LLM under data-residency constraints?
Home
Topics
Search
Library