Skip to main navigation Skip to search Skip to main content

Screen Detection From Egocentric Image Streams Leveraging Multi-View Vision Language Model

  • Xueshen Li
  • , Sen Shen
  • , Xinlong Hou
  • , Xinran Gao
  • , Ziyi Huang
  • , Steven J. Holiday
  • , Matthew R. Cribbet
  • , Susan W. White
  • , Edward Sazonov
  • , Yu Gan
  • Stevens Institute of Technology
  • Iowa State University
  • Columbia University
  • Arizona State University
  • University of Alabama

Research output: Contribution to journalArticlepeer-review

Abstract

Accurately monitoring the screen exposure of young children is important for research related to screen use such as childhood obesity, physical activity, and social interaction. Most existing studies rely upon self-report or manual measures from bulky wearable sensors, thus lacking efficiency and accuracy in capturing quantitative screen exposure data. In this work, we developed a novel screen detection framework that utilizes egocentric images from a wearable sensor, named the screen time tracker (STT), and a vision language model (VLM). In particular, we devised a multi-view VLM that takes multiple views from egocentric image streams and interprets screen exposure dynamically. We validated our approach by using a dataset of children’s free-living activities, demonstrating significant improvement over existing methods in conventional vision language models and object detection models. The combination of vision language model and lightweight hardware design provides a novel solution in screen detection for children. The proposed framework has great potential to benefit children’s behavioral study.

Original languageEnglish
Pages (from-to)4823-4834
Number of pages12
JournalIEEE Transactions on Multimedia
Volume28
DOIs
StatePublished - 2026

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Vision language model
  • egocentric image streams
  • wearable sensor

Fingerprint

Dive into the research topics of 'Screen Detection From Egocentric Image Streams Leveraging Multi-View Vision Language Model'. Together they form a unique fingerprint.

Cite this