← Back to all articles
arXiv cs.AIAugust 18, 2026

Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment

Excerpt

arXiv:2608.15693v1 Announce Type: new Abstract: Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real hardware and extract practical deployment guidelines from them. Following these guidelines, we deploy compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answerin