Zero-shot Active Visual Search (ZAVIS): Intelligent Object Search for Robotic Assistants

Jeongeun Park¹ Taerim Yoon¹ Jejoon Hong¹ Youngjae Yu² Matthew Pan³ Sungjoon Choi*¹
Korea University¹
Yonsei University²
Queens University³

To be appear in ICRA 2023

Abstract

In this paper, we focus on the problem of efficiently locating a target object described with free-form text using a mobile robot equipped with vision sensors (e.g., an RGBD camera). Conventional active visual search predefines a set of objects to search for, rendering these techniques restrictive in practice. To provide added flexibility in active visual searching, we propose a system where a user can enter target commands using free-form text; we call this system Zero-shot Active Visual Search (ZAVIS). ZAVIS detects and plans to search for a target object inputted by a user through a semantic grid map represented by static landmarks (e.g., desk or bed). For efficient planning of object search patterns, ZAVIS considers commonsense knowledge-based co-occurrence and predictive uncertainty while deciding which landmarks to visit first. We validate the proposed method with respect to SR (success rate) and SPL (success weighted by path length) in both simulated and real-world environments. The proposed method outperforms previous methods in terms of SPL in simulated scenarios with an average gap of 0.283. We further demonstrate ZAVIS with a Pioneer-3AT robot in real-world studies.

Proposed Method

Architecture

The overall procedure of ZAVIS. We present ZAVIS framework, object search based on free-form language. The framework consists of initial scanning, waypoint generation, navigation, and detection. The framework iteratively asks the human to compare the candidate patch set.

Detection

In order to search for objects given as free-form language, we also need to detect the objects whose label is not in the training dataset. The proposed open-set object detection module is illustrated on the above figure.

Zero-shot Active Visual Search (ZAVIS): Intelligent Object Search for Robotic Assistants

Jeongeun Park¹ Taerim Yoon¹ Jejoon Hong¹ Youngjae Yu² Matthew Pan³ Sungjoon Choi*¹
Korea University¹
Yonsei University²
Queens University³

Paper

Video

Code

To be appear in ICRA 2023

Abstract

Proposed Method

Architecture

Detection

Real-world Demonstration

Ablation studies on Real-world Environment

Real-world Environment on Various Tarets

Citation

Zero-shot Active Visual Search (ZAVIS): Intelligent Object Search for Robotic Assistants

Jeongeun Park1 Taerim Yoon1 Jejoon Hong1 Youngjae Yu2 Matthew Pan3 Sungjoon Choi*1 Korea University1 Yonsei University2 Queens University3

Paper

Video

Code

To be appear in ICRA 2023

Abstract

Proposed Method

Architecture

Detection

Real-world Demonstration

Ablation studies on Real-world Environment

Real-world Environment on Various Tarets

Citation

Jeongeun Park¹ Taerim Yoon¹ Jejoon Hong¹ Youngjae Yu² Matthew Pan³ Sungjoon Choi*¹
Korea University¹
Yonsei University²
Queens University³