TY - GEN
T1 - PopUp
T2 - 24th ACM International Conference on Intelligent User Interfaces, IUI 2019
AU - Song, Jean Y.
AU - Lemmer, Stephan J.
AU - Liu, Michael Xieyang
AU - Yan, Shiyan
AU - Kim, Juho
AU - Corso, Jason J.
AU - Lasecki, Walter S.
N1 - Publisher Copyright:
© 2019 Copyright held by the author(s). Publication rights licensed to ACM.
PY - 2019
Y1 - 2019
N2 - Collecting a sufficient amount of 3D training data for autonomous vehicles to handle rare, but critical, traffic events (e.g., collisions) may take decades of deployment. Abundant video data of such events from municipal traffic cameras and video sharing sites (e.g., YouTube) could provide a potential alternative, but generating realistic training data in the form of 3D video reconstructions is a challenging task beyond the current capabilities of computer vision. Crowdsourcing the annotation of necessary information could bridge this gap, but the level of accuracy required to obtain usable reconstructions makes this task nearly impossible for non-experts. In this paper, we propose a novel hybrid intelligence method that combines annotations from workers viewing different instances (video frames) of the same target (3D object), and uses particle filtering to aggregate responses. Our approach can leveraging temporal dependencies between video frames, enabling higher quality through more aggressive filtering. The proposed method results in a 33% reduction in the relative error of position estimation compared to a state-of-the-art baseline. Moreover, our method enables skipping (self-filtering) challenging annotations, reducing the total annotation time for hard-to-annotate frames by 16%. Our approach provides a generalizable means of aggregating more accurate crowd responses in settings where annotation is especially challenging or error-prone.
AB - Collecting a sufficient amount of 3D training data for autonomous vehicles to handle rare, but critical, traffic events (e.g., collisions) may take decades of deployment. Abundant video data of such events from municipal traffic cameras and video sharing sites (e.g., YouTube) could provide a potential alternative, but generating realistic training data in the form of 3D video reconstructions is a challenging task beyond the current capabilities of computer vision. Crowdsourcing the annotation of necessary information could bridge this gap, but the level of accuracy required to obtain usable reconstructions makes this task nearly impossible for non-experts. In this paper, we propose a novel hybrid intelligence method that combines annotations from workers viewing different instances (video frames) of the same target (3D object), and uses particle filtering to aggregate responses. Our approach can leveraging temporal dependencies between video frames, enabling higher quality through more aggressive filtering. The proposed method results in a 33% reduction in the relative error of position estimation compared to a state-of-the-art baseline. Moreover, our method enables skipping (self-filtering) challenging annotations, reducing the total annotation time for hard-to-annotate frames by 16%. Our approach provides a generalizable means of aggregating more accurate crowd responses in settings where annotation is especially challenging or error-prone.
KW - 3D Reconstruction
KW - Answer Aggregation
KW - Autonomous Vehicle
KW - Crowdsourcing
KW - Human Computation
KW - Particle Filter
UR - https://www.scopus.com/pages/publications/85065564793
UR - https://www.scopus.com/pages/publications/85065564793#tab=citedBy
U2 - 10.1145/3301275.3302305
DO - 10.1145/3301275.3302305
M3 - Conference contribution
AN - SCOPUS:85065564793
SN - 9781450362726
T3 - International Conference on Intelligent User Interfaces, Proceedings IUI
SP - 558
EP - 569
BT - Proceedings of the 24th International Conference on Intelligent User Interfaces
Y2 - 17 March 2019 through 20 March 2019
ER -