
Abstract Automated software engineering is rapidly transitioning from monolithic, cloud-bound Large Language Models (LLMs) to localized, optimized Small Language Models (SLMs). While networks exceeding 70-billion parameters represent the cognitive ceiling for complex reasoning, executing always-on code autocomplete within tight developer feedback loops requires sub-200 millisecond latencies. This paper provides a comprehensive analysis of the hardware-software co-design required to run code-synthesis SLMs natively on resource-constrained consumer hardware, specifically focusing on Unified Memory Architecture (UMA) APUs. We evaluate state-of-the-art architectures, detail spatial sequence formatting mechanics, diagnose memory bottlenecks under strict 8 GB system limits, and introduce optimization frameworks—including Zero-Copy DMA-BUF Inference, Adaptive Placeholder Completion (APC), and the SynConfRoute hybrid orchestration pipeline.
Small Language Models, Edge AI, Fill-in-the-Middle
Small Language Models, Edge AI, Fill-in-the-Middle
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
