TY - GEN
T1 - Semi-Supervised Masked Autoencoders
T2 - 35th Wireless and Optical Communications Conference, WOCC 2026
AU - Faysal, Atik
AU - Rostami, Mohammad
AU - Roshan, Reihaneh Gh
AU - Muralidhar, Nikhil
AU - Wang, Huaxia
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - We address the challenge of training Vision Transformers (ViTs) when labeled data is scarce but unlabeled data is abundant. We propose Semi-Supervised Masked Autoencoder (SSMAE), a framework that jointly optimizes masked image reconstruction and classification using both unlabeled and labeled samples with dynamically selected pseudo-labels. SSMAE introduces a validation-driven gating mechanism that activates pseudo-labeling only after the model achieves reliable, high-confidence predictions that are consistent across both weakly and strongly augmented views of the same image, reducing confirmation bias. On CIFAR-10 and CIFAR-100, SSMAE consistently outperforms supervised and self-supervised baselines, with the largest gains in low-label regimes (+9.24% over ViT-B on CIFAR-10 with 10% labels). Our results demonstrate that when pseudo-labels are introduced is as important as how they are generated for data-efficient transformer training. Codes are available at https://github.com/atik666/ssmae.
AB - We address the challenge of training Vision Transformers (ViTs) when labeled data is scarce but unlabeled data is abundant. We propose Semi-Supervised Masked Autoencoder (SSMAE), a framework that jointly optimizes masked image reconstruction and classification using both unlabeled and labeled samples with dynamically selected pseudo-labels. SSMAE introduces a validation-driven gating mechanism that activates pseudo-labeling only after the model achieves reliable, high-confidence predictions that are consistent across both weakly and strongly augmented views of the same image, reducing confirmation bias. On CIFAR-10 and CIFAR-100, SSMAE consistently outperforms supervised and self-supervised baselines, with the largest gains in low-label regimes (+9.24% over ViT-B on CIFAR-10 with 10% labels). Our results demonstrate that when pseudo-labels are introduced is as important as how they are generated for data-efficient transformer training. Codes are available at https://github.com/atik666/ssmae.
KW - image classification
KW - mask modeling
KW - pseudo-labeling
KW - representation learning
KW - semi-supervised learning
UR - https://www.scopus.com/pages/publications/105042812528
UR - https://www.scopus.com/pages/publications/105042812528#tab=citedBy
U2 - 10.1109/WOCC69802.2026.11556199
DO - 10.1109/WOCC69802.2026.11556199
M3 - Conference contribution
AN - SCOPUS:105042812528
T3 - 35th Wireless and Optical Communications Conference, WOCC 2026
BT - 35th Wireless and Optical Communications Conference, WOCC 2026
Y2 - 8 May 2026 through 9 May 2026
ER -