An LLM-Guided Agentic Optimization Toolbox for Edge-Deployable Neural Networks
| Autorzy | Nakip M.; Gibała R. |
|---|---|
| Tytuł | An LLM-Guided Agentic Optimization Toolbox for Edge-Deployable Neural Networks |
| Czasopismo | 35th ACM International Conference on Information and Knowledge Management |
| Rok | 2026 |
| Status | In Press |
| Abstrakt | <p><span style="-webkit-text-stroke-width:0px;background-color:rgb(255, 255, 255);color:rgb(0, 0, 0);display:inline !important;float:none;font-family:"Helvetica Neue", Helvetica, Arial, sans-serif;font-size:12px;font-style:normal;font-variant-caps:normal;font-variant-ligatures:normal;font-weight:400;letter-spacing:normal;orphans:2;text-align:left;text-decoration-color:initial;text-decoration-style:initial;text-decoration-thickness:initial;text-indent:0px;text-transform:none;white-space:normal;widows:2;word-spacing:0px;">Deploying neural networks on resource-constrained edge devices is limited by severe envelopes on storage, memory, compute, and energy. Manual compression through pruning, weight clustering, and quantization can reduce footprint substantially, yet selecting an appropriate strategy and its hyperparameters remains a search-heavy process. In this paper, we present an automated Edge AI Optimization Toolbox that closes the loop between weight-distribution profiling, user-defined multi-objective preferences, and Large Language Model (LLM)-guided strategy synthesis for Neural Networks targeting edge export. Given a pretrained model, a calibration and evaluation dataset, and constraints on minimum accuracy together with maximum model size and RAM, the toolbox profiles layer-wise weight statistics, proposes a compression strategy among pruning, clustering, or their combination, always applies int8 quantization, re-evaluates the candidate, and iterates until the preferences are met or a maximum iteration budget is exhausted, returning either an ideal or closest-matching model. The approach is aligned with Green AI goals of reducing unnecessary computational and memory footprints at deployment time. On occupancy estimation models spanning MLP, CNN, LSTM, and public pretrained architectures, the toolbox reduces model size by roughly 72%-92% while largely preserving predictive quality. Inference latency significantly improves for several models, whereas energy and carbon impacts remain architecture-dependent under the current operator scope. These results show that preference-constrained, LLM-guided compression is an effective and promising path toward edge-deployable and resource-aware neural networks.</span></p> |