Researchers have developed QScheduler, an adaptive algorithm designed to optimize on-device training for microcontrollers equipped with Neural Processing Units (NPUs). This method estimates gradients using only forward passes, making it suitable for devices without backpropagation capabilities and reducing memory usage. QScheduler dynamically adjusts the number of gradient samples based on training progress, eliminating the need for costly hyperparameter searches. Initial experiments on the STM32N6's Neural-ART NPU demonstrate that QScheduler achieves performance comparable to manually tuned fixed-q configurations for ResNet18 and MobileNetV2 models on datasets like EuroSAT and STL-10. AI
IMPACT Enables more efficient and accessible on-device AI training on resource-constrained hardware.
RANK_REASON The cluster contains a research paper detailing a new algorithm for on-device training. [lever_c_demoted from research: ic=1 ai=1.0]
- EuroSAT
- INT8 NPUs
- MobileNetV2
- QScheduler
- ResNet18
- STL-10
- STM32N6
- Victor Felipe Domingues do Amaral
- Zeroth-Order optimization
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →