
We demonstrate a system that replaces the default Kubernetes scheduler with an inference and topology-aware Reinforcement Learning policy. The system time-slices each GPU, consumes live Prometheus telemetry and places inference workloads across Far-Edge, Edge and Cloud tiers using an action masked agent (MaskablePPO, Maskable DQN, or DreamerV3) retrained and hot-swapped online without downtime. From a dashboard, attendees select a policy, submit workloads, shape network-function traffic and compare KPIs live on every device. Crucially, the placement behavior learned offline, is transferred to the live cluster: per-tier placement on real hardware track the offline simulation across all schedulers with TV < 0.05.
