VA- GKD: Variance- Aware Adaptive Group Knowledge Distillation for Long-Tailed Recognition
Keywords:
Long-tail distribution, Imbalanced data, Image classification, Adaptive variance, deep learning, Features, Synthetic Data, SABLC, Knowledge DistillationAbstract
Knowledge distillation (KD) in long-tailed recognition suffers from a
systematic bias: the teacher model’s output distribution is heavily skewed
towards the dominant head classes, causing the student model to inherit
and amplify head-class favoritism. Long-Tailed Knowledge Distillation
(LTKD) addresses this by partitioning classes into head, medium, and tail
groups and rebalancing group-level probabilities via a batch-mean scalar
correction. We identify two limitations of this approach: (i) the batchmean correction is a first-moment fix that ignores intra-batch variance,
leading to inconsistent per-sample corrections; and (ii) the intra-group replacement weight β is a static constant, blind to the temporal evolution of
the teacher’s bias during training. The proposed method Variance-Aware
Group Knowledge Distillation (VA-GKD), which augments LTKD with
(a) a sample-adaptive, variance-normalized cross-group scaling factor
derived from the per-batch distribution of group probabilities, and (b) a
cosine-scheduled curriculum on the intra-group replacement weight β(t)
that transitions from a teacher-proportional initialization to a uniform
emphasis encouraging tail-class learning. Experiments on a controlled
synthetic long-tailed dataset yield a statistically significant improvement
in tail-class accuracy over the baseline method












