Get Started
Home
Topics
Search
Library
Research questionHow should dependent LLM inference tasks be coordinated across edge servers when deadlines allow only limited extensions?A missed deadline for one dependent subtask can compromise an entire inference request. Distributed edge servers must therefore allocate work under tight latency limits while using only a constrained amount of deadline flexibility.
AI
Inference Optimization
Reinforcement Learning
Technology
Latest papersRecent research connected to this question, newest first.Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPOThe setting is collaborative mobile edge computing for LLM inference, with task dependencies and migration among MEC servers. The supplied evidence is simulation-based and compares a transformer-enhanced PPO framework with conventional PPO and heuristic approaches using task completion and overall system efficiency outcomes.research paper · Sep 3, 2026
Related questions
How can systems route each task to a suitable LLM endpoint under competing quality, cost, latency, and policy constraints?How can diffusion language models support reliable mobile-edge agents under tight latency and resource constraints?How can LLM orchestrators preserve continuous state when collaborating with non-language agents?How can enterprises consolidate heterogeneous internal workloads onto one self-hosted LLM under data-residency constraints?