TY - GEN
T1 - Break Out the Silverware
T2 - 28th International Conference on Pattern Recognition, ICPR 2026
AU - Levi Richter, Michaela
AU - Mirsky, Reuth
AU - Glickman, Oren
N1 - Publisher Copyright: © The Author(s), under exclusive license to Springer Nature Switzerland AG 2027.
PY - 2027
Y1 - 2027
N2 - “Bring me a plate.” For domestic service robots, such a simple command exposes a fundamental challenge: inferring where everyday objects are stored when they are not directly visible, such as inside drawers, cabinets, or closets. While recent advances in vision and manipulation have improved robotic perception, robots still lack the commonsense reasoning needed to predict plausible storage locations. We introduce the Stored Household Item Challenge, a benchmark task for evaluating a robot’s ability to infer likely storage locations given a household kitchen scene and a queried item. The benchmark comprises two datasets: (1) a real-world evaluation set of 100 item–image pairs with human-annotated ground truth from participants’ kitchens, and (2) a development set of 6,500 item–image pairs with polygon-level storage annotations on public kitchen images. To address this task, we propose NOAM (Non-visible Object Allocation Model), a hybrid vision–language pipeline that converts visual input into structured natural-language descriptions of spatial context and visible containers, and then prompts a large language model to infer the most likely hidden storage location. We evaluate NOAM against random baselines, vision–language pipelines, state-of-the-art multimodal models, and human performance. NOAM substantially improves prediction accuracy and approaches human-level results, demonstrating the value of structured vision–language reasoning for cognitively capable service robots in domestic environments.
AB - “Bring me a plate.” For domestic service robots, such a simple command exposes a fundamental challenge: inferring where everyday objects are stored when they are not directly visible, such as inside drawers, cabinets, or closets. While recent advances in vision and manipulation have improved robotic perception, robots still lack the commonsense reasoning needed to predict plausible storage locations. We introduce the Stored Household Item Challenge, a benchmark task for evaluating a robot’s ability to infer likely storage locations given a household kitchen scene and a queried item. The benchmark comprises two datasets: (1) a real-world evaluation set of 100 item–image pairs with human-annotated ground truth from participants’ kitchens, and (2) a development set of 6,500 item–image pairs with polygon-level storage annotations on public kitchen images. To address this task, we propose NOAM (Non-visible Object Allocation Model), a hybrid vision–language pipeline that converts visual input into structured natural-language descriptions of spatial context and visible containers, and then prompts a large language model to infer the most likely hidden storage location. We evaluate NOAM against random baselines, vision–language pipelines, state-of-the-art multimodal models, and human performance. NOAM substantially improves prediction accuracy and approaches human-level results, demonstrating the value of structured vision–language reasoning for cognitively capable service robots in domestic environments.
KW - Benchmark Datasets
KW - Commonsense Reasoning
KW - Domestic Robotics
KW - Large Language Models
KW - Semantic Understanding
UR - https://www.scopus.com/pages/publications/105046971664
U2 - 10.1007/978-3-032-31930-2_22
DO - 10.1007/978-3-032-31930-2_22
M3 - Conference contribution
SN - 9783032319296
T3 - Lecture Notes in Computer Science
SP - 322
EP - 336
BT - Pattern Recognition - 28th International Conference, ICPR 2026, Proceedings
A2 - De Marsico, Maria
A2 - Ho, Tin Kam
A2 - Jurie, Frederic
A2 - Liu, Cheng-Lin
A2 - Lopresti, Daniel
A2 - Nyström, Ingela
A2 - Ogier, Jean-Marc
A2 - Ross, Arun
A2 - Wang, Liang
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 17 August 2026 through 22 August 2026
ER -