Skip to main navigation Skip to search Skip to main content

Automatic Image Labeling Framework Based on Multiple Vision-Language Models for Automation Applications

  • Xiaowei Su
  • , Hong Chen
  • , Kai Liu
  • , Yudong Yao
  • , Endai Huang
  • Ningbo University
  • City University of Hong Kong

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Agricultural automation and robotics are increasingly transforming modern livestock farming, with computer vision and artificial intelligence algorithms playing a critical role. However, vision-based robotic systems rely heavily on large, well-annotated datasets to train reliable AI models, while manual labeling remains a labor-intensive and time-consuming process. To address this challenge, we propose an automatic image labeling framework based on multiple vision-language models. The framework integrates target point detection, mask generation and fusion, and a mask filtering and voting mechanism to produce high-quality instance masks across diverse livestock species and imaging conditions. Experimental results demonstrate that our method outperforms existing approaches such as Grounded-SAM and SEEM. By enabling efficient dataset creation, the proposed framework supports the development of vision-based perception systems essential for agricultural robotics and intelligent automation.

Original languageEnglish
Title of host publicationProceedings - 2025 10th International Conference on Robotics and Automation Engineering, ICRAE 2025
Pages261-265
Number of pages5
ISBN (Electronic)9798331550257
DOIs
StatePublished - 2025
Event10th International Conference on Robotics and Automation Engineering, ICRAE 2025 - Haikou, China
Duration: 14 Nov 202516 Nov 2025

Publication series

NameProceedings - 2025 10th International Conference on Robotics and Automation Engineering, ICRAE 2025

Conference

Conference10th International Conference on Robotics and Automation Engineering, ICRAE 2025
Country/TerritoryChina
CityHaikou
Period14/11/2516/11/25

Keywords

  • Animal Monitoring
  • Deep Learning
  • Instance Segmentation
  • Pseudo-labeling
  • Zero-shot Labeling

Fingerprint

Dive into the research topics of 'Automatic Image Labeling Framework Based on Multiple Vision-Language Models for Automation Applications'. Together they form a unique fingerprint.

Cite this