Robot Learning to Communicate through Projected Visual Abstractions

Agentic AI
Published: arXiv: 2607.22434v1
Authors

Danyang Yan Boyuan Wang Jiaxun Liu Boyuan Chen

Abstract

Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through such projected visual abstractions requires reasoning not only about bodily motion but also about how that motion is transformed into an external representation perceived by an observer. Among these abstractions, shadows provide a particularly compelling example because they emerge directly from the robot's embodiment while remaining visually distinct from the body itself. Here, we present a robotic system capable of dynamic shadow expression using a 21-degree-of-freedom dexterous hand with compliant soft skin and a learned shadow self-model. The soft-skinned embodiment reduces light leakage to produce visually continuous silhouettes, while the differentiable self-model learns the mapping between hand configurations and projected shadow appearance through task-agnostic self-exploration. Given a target shadow image or video, the robot optimizes its hand configurations through gradient-based search over 1 the learned self-model and refines the solution through collision-aware simulation to obtain physically feasible motions. For dynamic shadow performance, we further introduce expressive-region objectives, temporal smoothness regularization, and keyframe-based optimization to preserve visually important motion cues while reducing optimization complexity. We demonstrate robotic shadow expression across sign-language gestures, hand-shadow puppetry, and animal motion imitation in both simulation and physical experiments. These results establish a framework for enabling robots to manipulate projected visual abstractions of themselves for communication and visual storytelling.

Paper Summary

Problem
The main problem addressed in this paper is that robots are limited in their ability to communicate through projected visual abstractions, such as shadows, silhouettes, and reflections. While humans use these abstractions to convey meaning and communicate, robots are confined to expressing themselves through their physical morphology. The researchers aim to enable robots to communicate through projected visual abstractions, starting with shadows, which emerge directly from a robot's embodiment.
Key Innovation
The key innovation of this work is a robotic system capable of dynamic shadow expression using a 21-degree-of-freedom dexterous hand with compliant soft skin and a learned shadow self-model. The soft-skinned embodiment reduces light leakage to produce visually continuous silhouettes, while the differentiable self-model learns the mapping between hand configurations and projected shadow appearance through task-agnostic self-exploration.
Practical Impact
This research has practical implications for robotics and human-robot interaction. By enabling robots to communicate through projected visual abstractions, robots can better express themselves and convey meaning to humans. This can lead to more effective human-robot collaboration and communication in various applications, such as education, entertainment, and assistive technology. Additionally, this research can inspire new forms of human-robot interaction, such as using shadows and silhouettes to convey emotions and intentions.
Analogy / Intuitive Explanation
Imagine you're watching a puppet show, where the puppeteer uses shadows and silhouettes to create a story. The puppeteer's hand movements and the lighting create a dynamic and expressive visual representation that conveys meaning to the audience. Similarly, the robotic system developed in this research uses a dexterous hand and soft skin to create dynamic shadows that can be used to communicate with humans. The system learns to map hand configurations to projected shadow appearance, allowing it to express itself in a more human-like way.
Paper Information
Categories:
cs.RO cs.AI
Published Date:

arXiv ID:

2607.22434v1

Quick Actions