TY - GEN
T1 - Energy-Efficient Hybrid AI Tutor for Learning Varieties of English
AU - You, Cunqian
AU - Wei, Miao
AU - Wang, Xiaojun
AU - Lu, Huijuan
AU - Yao, Yudong
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Energy-aware design is becoming a practical constraint for speech-first language tutors, especially when learners expect fluent interaction across multiple varieties of English (AmE, BrE, Indian English, AusE, and CanE). We compare edge-only, cloud-only, and hybrid edge-cloud tutoring under a fully specified simulation that includes an on-device ASR profile, a compressed local tutor model, cloud GPU allocation with PUE, transport energy, and repeated trials. The revised evaluation reports not only energy and latency but also normalized learning gain ΔS and learning efficiency LE=ΔS/Etotal. Across 150 simulated 5-minute sessions, the hybrid design with caching uses 0.37+/-0.04 Wh, offloads 15.6+/- 4.8% of turns, preserves dialect-style accuracy within about one point of the cloud baseline, and improves LE over cloud-only by 46% when LE is reported as gain per Wh. These results suggest that uncertainty-gated hybrid inference can provide a practical energy- accuracy trade-off for dialect-aware tutoring, while the present findings should be interpreted as simulation-based rather than as evidence from physical deployment or learner trials.
AB - Energy-aware design is becoming a practical constraint for speech-first language tutors, especially when learners expect fluent interaction across multiple varieties of English (AmE, BrE, Indian English, AusE, and CanE). We compare edge-only, cloud-only, and hybrid edge-cloud tutoring under a fully specified simulation that includes an on-device ASR profile, a compressed local tutor model, cloud GPU allocation with PUE, transport energy, and repeated trials. The revised evaluation reports not only energy and latency but also normalized learning gain ΔS and learning efficiency LE=ΔS/Etotal. Across 150 simulated 5-minute sessions, the hybrid design with caching uses 0.37+/-0.04 Wh, offloads 15.6+/- 4.8% of turns, preserves dialect-style accuracy within about one point of the cloud baseline, and improves LE over cloud-only by 46% when LE is reported as gain per Wh. These results suggest that uncertainty-gated hybrid inference can provide a practical energy- accuracy trade-off for dialect-aware tutoring, while the present findings should be interpreted as simulation-based rather than as evidence from physical deployment or learner trials.
KW - cache-assisted inference
KW - dialect-aware tutoring
KW - edge-cloud collaboration
KW - energy-efficient AI
KW - model compression
KW - speech interfaces
KW - uncertainty-based offloading
UR - https://www.scopus.com/pages/publications/105042709975
UR - https://www.scopus.com/pages/publications/105042709975#tab=citedBy
U2 - 10.1109/WOCC69802.2026.11556181
DO - 10.1109/WOCC69802.2026.11556181
M3 - Conference contribution
AN - SCOPUS:105042709975
T3 - 35th Wireless and Optical Communications Conference, WOCC 2026
BT - 35th Wireless and Optical Communications Conference, WOCC 2026
T2 - 35th Wireless and Optical Communications Conference, WOCC 2026
Y2 - 8 May 2026 through 9 May 2026
ER -