02527nas a2200097 4500000000100000008004100001100001500042700001900057245008300076520227000159 2026 d1 aMert Nakip1 aRafał Gibała00aAn LLM-Guided Agentic Optimization Toolbox for Edge-Deployable Neural Networks3 a
Deploying neural networks on resource-constrained edge devices is limited by severe envelopes on storage, memory, compute, and energy. Manual compression through pruning, weight clustering, and quantization can reduce footprint substantially, yet selecting an appropriate strategy and its hyperparameters remains a search-heavy process. In this paper, we present an automated Edge AI Optimization Toolbox that closes the loop between weight-distribution profiling, user-defined multi-objective preferences, and Large Language Model (LLM)-guided strategy synthesis for Neural Networks targeting edge export. Given a pretrained model, a calibration and evaluation dataset, and constraints on minimum accuracy together with maximum model size and RAM, the toolbox profiles layer-wise weight statistics, proposes a compression strategy among pruning, clustering, or their combination, always applies int8 quantization, re-evaluates the candidate, and iterates until the preferences are met or a maximum iteration budget is exhausted, returning either an ideal or closest-matching model. The approach is aligned with Green AI goals of reducing unnecessary computational and memory footprints at deployment time. On occupancy estimation models spanning MLP, CNN, LSTM, and public pretrained architectures, the toolbox reduces model size by roughly 72%-92% while largely preserving predictive quality. Inference latency significantly improves for several models, whereas energy and carbon impacts remain architecture-dependent under the current operator scope. These results show that preference-constrained, LLM-guided compression is an effective and promising path toward edge-deployable and resource-aware neural networks.