TY - GEN
T1 - Open-Vocabulary Semantic Segmentation Using Test-Time Distillation
AU - Zabari, Nir
AU - Hoshen, Yedid
N1 - Publisher Copyright: © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
PY - 2023
Y1 - 2023
N2 - Semantic segmentation is a key computer vision task that has been actively researched for decades. In recent years, supervised methods have reached unprecedented accuracy; however, obtaining pixel-level annotation is very time-consuming and expensive. In this paper, we propose a novel open-vocabulary approach to creating semantic segmentation masks, without the need for training segmentation networks or seeing any segmentation masks. At test time, our method takes as input the image-level labels of the categories present in the image. We utilize a vision-language embedding model to create a rough segmentation map for each class via model interpretability methods and refine the maps using a test-time augmentation technique. The output of this stage provides pixel-level pseudo-labels, which are utilized by single-image segmentation techniques to obtain high-quality output segmentations. Our method is shown quantitatively and qualitatively to outperform methods that use a similar amount of supervision, and to be competitive with weakly-supervised semantic-segmentation techniques.
AB - Semantic segmentation is a key computer vision task that has been actively researched for decades. In recent years, supervised methods have reached unprecedented accuracy; however, obtaining pixel-level annotation is very time-consuming and expensive. In this paper, we propose a novel open-vocabulary approach to creating semantic segmentation masks, without the need for training segmentation networks or seeing any segmentation masks. At test time, our method takes as input the image-level labels of the categories present in the image. We utilize a vision-language embedding model to create a rough segmentation map for each class via model interpretability methods and refine the maps using a test-time augmentation technique. The output of this stage provides pixel-level pseudo-labels, which are utilized by single-image segmentation techniques to obtain high-quality output segmentations. Our method is shown quantitatively and qualitatively to outperform methods that use a similar amount of supervision, and to be competitive with weakly-supervised semantic-segmentation techniques.
KW - Language-based segmentation
KW - Open-vocabulary segmentation
KW - Semantic segmentation
UR - https://www.scopus.com/pages/publications/85151063978
U2 - 10.1007/978-3-031-25063-7_4
DO - 10.1007/978-3-031-25063-7_4
M3 - Conference contribution
SN - 9783031250620
T3 - Lecture Notes in Computer Science
SP - 56
EP - 72
BT - Computer Vision – ECCV 2022 Workshops, Proceedings
A2 - Karlinsky, Leonid
A2 - Michaeli, Tomer
A2 - Nishino, Ko
PB - Springer Science and Business Media Deutschland GmbH
T2 - Workshops held at the 17th European Conference on Computer Vision, ECCV 2022
Y2 - 23 October 2022 through 27 October 2022
ER -