Research questionHow can complex neural-network graphs be mapped across heterogeneous SoCs to balance inference latency and throughput?Complex neural-network graphs expose parallel operators, but executing them across unlike on-chip processors introduces dependencies and communication costs. Choices that improve pipeline throughput can worsen single-inference latency or energy efficiency.