Abstract:This paper proposes a Hierarchical Illumination-Token Fusion Network to explicitly bridge global exposure distribution and local restoration for image exposure correction. The method employs a hierarchical token fusion mechanism with contrastive learning to enforce global illumination consistency. A cross-modality multi-scale fusion module adaptively merges complementary features under global modulation, while a frequency-guided dynamic attention module leverages frequency priors to recover high-frequency structures. Experiments on MSEC and SICE datasets show that the network achieves average peak signal-to-noise ratio values of 23.72 dB and 21.92 dB, respectively. This demonstrates outstanding performance while utilizing significantly fewer parameters than leading methods, establishing a new efficiency bench-mark.