RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction

Zifan Wang*1,2, Ziang Ren*2,3, Pengyang Shi*1,2, Zirui Wang4, Chenghuai Lin2, Tianze Wang2, Zekun Qi1,2, Liangliang Zhao4, He Wang2,5, Li Yi✉1,2,6
1Tsinghua University, 2Galbot, 3Beijing Institute of Technology, 4Harbin Institute of Technology, 5Peking University, 6Shanghai Qi Zhi Institute
*Equal contribution.  Corresponding author.
RoboGesture teaser

RoboGesture generates expressive, semantically-aligned, and safety-aware co-speech gestures for humanoid robots from streaming speech, enabling robots to listen, respond, and gesture coherently in real-time social interaction.

Co-Speech Gesture Generation

Social Interaction

Abstract

Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot interaction. However, this task faces three critical barriers: the scarcity of semantically rich datasets, the "modality eclipse" where models ignore audio cues in favor of kinematic inertia, and the sim-to-real gap regarding physical safety. We propose RoboGesture, a robot-centric framework that co-designs data, modeling, and control to power a complete interactive human–humanoid system in which the robot listens, responds, and gestures in real time.

We first establish the RoboGesture dataset featuring over 300 gesture categories and develop an automated pipeline to synthesize large-scale collision-free, robot-specific audio-motion pairs. Our architecture features a Hierarchical Semantic-Acoustic Aligner that extracts multi-granular prosodic and semantic cues directly from raw audio tokens. These cues drive a Streaming Conditional Motion Generator based on a diffusion transformer with conditional flow matching. To ensure high responsiveness, we introduce Anti-Inertia CFG Masking, which prevents the model from collapsing into repetitive historical patterns by compelling it to proactively mine control signals from the audio modality. Finally, an MPC-based safety filter ensures real-time, collision-free execution on physical hardware. Experiments on a Unitree G1 humanoid demonstrate that RoboGesture generates safer, more rhythmic, and more semantically appropriate responses compared to state-of-the-art baselines.

Video

Method

RoboGesture is a robot-centric, data–model co-designed framework that operates in a streaming-to-streaming paradigm: it processes incoming audio chunks and generates corresponding motions in real time. For the physical embodiment, we employ the Unitree G1 humanoid integrated with BrainCo dexterous hands, with a kinematic state space of 41 DoFs (17 for the upper torso and arms, 24 for the dexterous hands). The framework consists of three core modules: (i) a Hierarchical Semantic-Acoustic Aligner, (ii) a Streaming Conditional Motion Generator, and (iii) an MPC-based kinematic safety filter.

RoboGesture pipeline.

Overview of RoboGesture. (a) The Semantic-Acoustic Aligner decouples multi-granular audio cues from streaming tokens. (b) The Motion Generator synthesizes motion via Conditional Flow Matching, jointly conditioned on hierarchical audio features and historical motion. (c) At inference, the pipeline integrates with upstream Speech-LLMs and a safety filter to drive physical robots in real time.

Hierarchical Semantic-Acoustic Aligner

Text-based methods inevitably lose vital prosodic information such as intonation and emphasis — the word "Really" conveys skepticism with a rising intonation but confirmation with a falling one. To capture these nuances, we tokenize streaming audio with the Mimi codec and process the multi-granular tokens through a transformer-based aligner trained via multi-task auxiliary learning.

Shallow layers capture rapid acoustic energy transients and are routed to a beat head for rhythmic supervision, ensuring motion stays synchronized with acoustic onsets. Deep layers distill macroscopic semantics through a 300-class classification task powered by the RoboGesture dataset, acting as an information bottleneck. The resulting low-level rhythmic pulses and high-level semantic labels work in tandem to guide the downstream motion generator.

Streaming Conditional Motion Generator

Operating directly in the continuous kinematic state space, a Diffusion Transformer (DiT) with Conditional Flow Matching predicts the velocity field of future motion chunks for low-latency, high-fidelity streaming. To overcome the modality eclipse — where the model lazily relies on past kinematic inertia rather than audio — we adopt a two-stage training strategy with Anti-Inertia CFG Masking that randomly masks historical context, compelling the model to proactively mine control signals from audio. Inside the DiT, Cross-Attention provides frame-level alignment while FiLM injects global semantic modulation.

MPC-based Safety Filter

For real-time deployment, generated motion passes through an online MPC filter as a final safety guard, formulated as convex quadratic programming with velocity regularization, tracking-error minimization, and collision-avoidance constraints. The filter adds only 5.6 ms per frame, comfortably exceeding the 30 Hz robot control rate, and is also used offline to sanitize the entire training corpus.

RoboGesture Dataset

We build RoboGesture, a high-quality semantic gesture dataset for humanoid robots comprising over 300 gesture categories. The class list is derived from SeG and EgoGesture templates and expanded through user questionnaires. Each gesture is recorded with a marker-based motion-capture system, manually annotated with detailed motion descriptions, retargeted and optimized, then replayed on the physical robot to verify reachability and control fidelity.

Building on this, an automatic semi-synthetic pipeline generates millions of audio–motion pairs through five stages: (1) Scenario Generation with LLMs, (2) Gesture Tagging to select and place semantically appropriate gestures, (3) Multimodal Synthesis of emotion-aware TTS audio with word-level timestamps and rhythmic beat gestures, (4) Temporal Blending where semantic gestures begin 0.4s before their spoken keywords to mimic human anticipation, and (5) Safety Optimization via collision-avoidance refinement.

RoboGesture dataset.

Visualization of the RoboGesture Dataset. RoboGesture refines SeG by correcting motion-penetration issues, augments EgoGesture with full-body motion, and incorporates newly collected gestures. Each sample is annotated with detailed semantic labels.

Experiments

Co-Speech Gesture Generation

On the BEAT and SemanticBEAT benchmarks, RoboGesture achieves state-of-the-art FGD, BC, and MSE, indicating superior alignment with the ground-truth distribution and accurate structural reconstruction, while ensuring strict physical plausibility with a near-zero collision rate (0.88% / 0.13%).

Method BEAT SemanticBEAT
FGD ↓BC ↑MSE ↓DIV ↑Col. ↓ BC ↑DIV ↑Col. ↓
LivelySpeaker (ICCV'23) 3.03500.18110.18610.148521.41 0.28520.138413.36
DiffSHEG (CVPR'24) 2.23160.18510.17530.12180.85 0.28590.12091.20
SemTalk (ICCV'25) 7.93280.18280.39310.178152.82 0.29130.144042.65
Semantic Gesticulator (SIGGRAPH'24) 3.01470.17710.23750.281813.11 0.29060.283113.84
Ours 0.84520.18660.13470.20750.88 0.29500.20410.13

Quantitative comparison on BEAT and SemanticBEAT. Bold = best, underline = second-best.

Qualitative comparison.

Qualitative comparison of generated co-speech gestures. Left: baselines occasionally suffer severe self-penetration, whereas our model strictly avoids self-collisions. Middle: given the cue "OK", our model synthesizes the precise fine-grained hand gesture. Right: for emphatic speech ("seize the chance"), our model generates highly expressive body language such as confidently patting the chest.

Simulation Comparison with Baselines

Side-by-side comparisons of generated gestures highlight the superior semantic alignment, hand articulation, and physical safety of our method against state-of-the-art baselines.

Real-world Social Interaction

We instantiate a complete interactive human–humanoid system in which the robot listens, responds, and gestures in real time. The deployed system comprises a language interaction module (ASR, a LoRA-tuned LLM, and TTS), our streaming speech-to-motion module, and a robot execution module. The streaming motion generator sustains ≈120 FPS, with total system latency dominated by the upstream speech pipeline rather than our motion model. The carousel above shows real-world co-speech gesture generation and social interaction on a Unitree G1 humanoid.

BibTeX

@inproceedings{wang2026robogesture,
  title={RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction},
  author={Wang, Zifan and Ren, Ziang and Shi, Pengyang and Wang, Zirui and Lin, Chenghuai
         and Wang, Tianze and Qi, Zekun and Zhao, Liangliang and Wang, He and Yi, Li},
  booktitle={Proceedings of the European Conference on Computer Vision (ECCV)},
  year={2026}
}