文章背景与核心概要
在移动设备上执行深度神经网络(DNN)推理时,往往会因硬件资源受限而面临高延迟和高能耗的瓶颈。传统研究通常聚焦于针对计算单元的动态电压与频率调整(DVFS),而本文则创新性地将内存频率纳入关键变量。通过将内存频率、计算频率与通信资源进行联合优化,作者为实现高能效推理提供了一个稳健的理论框架。
该研究针对本地推理推导出了接近最优的闭式解,并针对基于边缘端的数据传输功率给出了最优解,同时辅以适用于实际部署的低复杂度启发式算法。仿真结果表明,该方法能够实现接近理论最优的性能(误差在 2.5% 以内),与基准方法相比,可将设备的能耗降低高达 10.4%。
Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference
Authors: Yunchu Han, Zhaojun Nan, Sheng Zhou, Zhisheng Niu
Date: August 14, 2026
Subject: Artificial Intelligence (cs.AI)
DOI: 10.48550/arXiv.2608.13863
Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference
Authors: Yunchu Han, Zhaojun Nan, Sheng Zhou, Zhisheng Niu
Date: August 14, 2026
Subject: Artificial Intelligence (cs.AI)
DOI: 10.48550/arXiv.2608.13863
Summary
深度神经网络(DNN)在移动设备上的推理经常由于硬件限制而受到高延迟和过高能耗的阻碍。虽然现有研究通常侧重于计算的动态电压和频率缩放(DVFS),但本文引入了一种将内存频率作为关键变量的新方法。通过将内存频率和计算频率与通信资源进行联合优化,作者为高能效推理提供了一个稳健的框架。该研究为本地推理提供了接近最优的闭式解,为基于边缘的传输功率提供了最优解,并辅以用于实际实现的低复杂度启发式算法。
Summary
Deep Neural Network (DNN) inference on mobile devices is frequently hindered by high latency and excessive energy consumption due to hardware constraints. While existing research typically focuses on Dynamic Voltage and Frequency Scaling (DVFS) for computing, this paper introduces a novel approach that incorporates memory frequency as a critical variable. By jointly optimizing memory and computing frequencies alongside communication resources, the authors provide a robust framework for energy-efficient inference. The study offers a near-optimal closed-form solution for local inference and an optimal solution for edge-based transmission power, supported by a low-complexity heuristic algorithm for practical implementation.
Key Contributions
- 整体优化模型: 与孤立计算频率的传统方法不同,本研究对内存频率和计算频率对总推理时间的影响进行了建模。
- 数学框架: 作者构建了一个优化问题,旨在最小化设备总能耗,同时严格遵守延迟截止期限。
- 高效算法:
- 本地推理: 通过凸优化推导出的接近最优的闭式解。
- 边缘推理: 在固定带宽下针对传输功率的最优闭式解。
- 启发式方法: 专为现实世界部署设计的多项式时间复杂度的低复杂度算法。
- 性能验证: 仿真结果表明,所提出的方法实现了接近最优的性能(在理论最优值的 2.5% 以内),与基准方法相比,设备的能耗降低了高达 10.4%。
Key Contributions
- Holistic Optimization Model: Unlike traditional methods that isolate computing frequency, this research models the impact of both memory and computing frequencies on total inference time.
- Mathematical Framework: The authors formulate an optimization problem aimed at minimizing total device energy consumption while strictly adhering to latency deadlines.
- Efficient Algorithms:
- Local Inference: A near-optimal closed-form solution derived via convex optimization.
- Edge Inference: An optimal closed-form solution for transmission power given a fixed bandwidth.
- Heuristic Approach: A low-complexity algorithm designed for real-world deployment with polynomial time complexity.
- Performance Validation: Simulation results demonstrate that the proposed method achieves near-optimal performance (within 2.5% of the theoretical optimum) and reduces device energy consumption by up to 10.4% compared to baseline methods.
Accessing the Paper
Accessing the Paper
Citation & Metadata
- 引用方式: arXiv:2608.13863 [cs.AI]
- 参考文献: NASA ADS | Google Scholar | Semantic Scholar
Citation & Metadata
- Cite as: arXiv:2608.13863 [cs.AI]
- References: NASA ADS | Google Scholar | Semantic Scholar