TY - GEN
T1 - Automatic Image Labeling Framework Based on Multiple Vision-Language Models for Automation Applications
AU - Su, Xiaowei
AU - Chen, Hong
AU - Liu, Kai
AU - Yao, Yudong
AU - Huang, Endai
N1 - Publisher Copyright:
©2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Agricultural automation and robotics are increasingly transforming modern livestock farming, with computer vision and artificial intelligence algorithms playing a critical role. However, vision-based robotic systems rely heavily on large, well-annotated datasets to train reliable AI models, while manual labeling remains a labor-intensive and time-consuming process. To address this challenge, we propose an automatic image labeling framework based on multiple vision-language models. The framework integrates target point detection, mask generation and fusion, and a mask filtering and voting mechanism to produce high-quality instance masks across diverse livestock species and imaging conditions. Experimental results demonstrate that our method outperforms existing approaches such as Grounded-SAM and SEEM. By enabling efficient dataset creation, the proposed framework supports the development of vision-based perception systems essential for agricultural robotics and intelligent automation.
AB - Agricultural automation and robotics are increasingly transforming modern livestock farming, with computer vision and artificial intelligence algorithms playing a critical role. However, vision-based robotic systems rely heavily on large, well-annotated datasets to train reliable AI models, while manual labeling remains a labor-intensive and time-consuming process. To address this challenge, we propose an automatic image labeling framework based on multiple vision-language models. The framework integrates target point detection, mask generation and fusion, and a mask filtering and voting mechanism to produce high-quality instance masks across diverse livestock species and imaging conditions. Experimental results demonstrate that our method outperforms existing approaches such as Grounded-SAM and SEEM. By enabling efficient dataset creation, the proposed framework supports the development of vision-based perception systems essential for agricultural robotics and intelligent automation.
KW - Animal Monitoring
KW - Deep Learning
KW - Instance Segmentation
KW - Pseudo-labeling
KW - Zero-shot Labeling
UR - https://www.scopus.com/pages/publications/105034394791
UR - https://www.scopus.com/pages/publications/105034394791#tab=citedBy
U2 - 10.1109/ICRAE67496.2025.00048
DO - 10.1109/ICRAE67496.2025.00048
M3 - Conference contribution
AN - SCOPUS:105034394791
T3 - Proceedings - 2025 10th International Conference on Robotics and Automation Engineering, ICRAE 2025
SP - 261
EP - 265
BT - Proceedings - 2025 10th International Conference on Robotics and Automation Engineering, ICRAE 2025
T2 - 10th International Conference on Robotics and Automation Engineering, ICRAE 2025
Y2 - 14 November 2025 through 16 November 2025
ER -