TY - GEN
T1 - Imitation-Inspired Semantic-Guided Distillation for User-Conditioned Memorability Prediction
AU - Ghosh, Indrajeet
AU - Anwar, Mohammad Saeid
AU - Jayarajah, Kasthuri
AU - Roy, Nirmalya
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Image memorability (IM) estimation typically relies on learning generic semantic features from large-scale datasets; however, memorability is intrinsically individual-dependent and personalized visual viewing behavior, but overlooking this variability can undermine model performance in downstream cognitive and human-machine interaction tasks, leading towards suboptimal performance. To address this, we propose MemGaze, a unified framework that integrates knowledge distillation (KD) and imitation learning (IL) to jointly learn generic and personalized salient representations for memorability estimation by leveraging both image content and gaze-derived heatmaps. MemGaze employs a teacher network built upon a pretrained ResNet-50 backbone, followed by an encoder-decoder architecture coupled with spatial and channel attention mechanisms to generate generic saliency-aware memorability maps. A lightweight, attention-guided student encoder-decoder is then optimized through a composite imitation-guided distillation process, where knowledge is distilled from the teacher while simultaneously imitating user-specific gaze fixation heatmaps. Through this joint training process, the student network learns to produce personalized memorability estimates while achieving substantial reductions in computational complexity. We validate MemGaze on two public IM estimation datasets (LaMem and SUN) and a inhouse WoM dataset comprising of 45 participants (UMBC IRB #670) engaged in visual search and navigation tasks that reflect individualized visual attention patterns. MemGaze outperforms nine state-of-the-art IM models by capturing coarse-to-fine saliency and adapting to individual attention, achieving ≈6% improvement in memorability prediction.
AB - Image memorability (IM) estimation typically relies on learning generic semantic features from large-scale datasets; however, memorability is intrinsically individual-dependent and personalized visual viewing behavior, but overlooking this variability can undermine model performance in downstream cognitive and human-machine interaction tasks, leading towards suboptimal performance. To address this, we propose MemGaze, a unified framework that integrates knowledge distillation (KD) and imitation learning (IL) to jointly learn generic and personalized salient representations for memorability estimation by leveraging both image content and gaze-derived heatmaps. MemGaze employs a teacher network built upon a pretrained ResNet-50 backbone, followed by an encoder-decoder architecture coupled with spatial and channel attention mechanisms to generate generic saliency-aware memorability maps. A lightweight, attention-guided student encoder-decoder is then optimized through a composite imitation-guided distillation process, where knowledge is distilled from the teacher while simultaneously imitating user-specific gaze fixation heatmaps. Through this joint training process, the student network learns to produce personalized memorability estimates while achieving substantial reductions in computational complexity. We validate MemGaze on two public IM estimation datasets (LaMem and SUN) and a inhouse WoM dataset comprising of 45 participants (UMBC IRB #670) engaged in visual search and navigation tasks that reflect individualized visual attention patterns. MemGaze outperforms nine state-of-the-art IM models by capturing coarse-to-fine saliency and adapting to individual attention, achieving ≈6% improvement in memorability prediction.
KW - Behavioral Cloning
KW - Image Memorability
KW - Knowledge Distillation
KW - Semantic Attention Modeling
UR - https://www.scopus.com/pages/publications/105035084144
UR - https://www.scopus.com/pages/publications/105035084144#tab=citedBy
U2 - 10.1109/ICDM65498.2025.00035
DO - 10.1109/ICDM65498.2025.00035
M3 - Conference contribution
AN - SCOPUS:105035084144
T3 - Proceedings - IEEE International Conference on Data Mining, ICDM
SP - 277
EP - 286
BT - Proceedings - 25th IEEE International Conference on Data Mining, ICDM 2025
A2 - Ding, Wei
A2 - Vreeken, Jilles
A2 - Lu, Chang-Tien
A2 - Gunopulos, Dimitrios
A2 - Wu, Xindong
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 25th IEEE International Conference on Data Mining, ICDM 2025
Y2 - 12 November 2025 through 15 November 2025
ER -