building the most expressive robot harness so far
If you want to skip the entire article and jump into the codebases, here you go:
- reachy-motion-generator: the models, training code, inference server, and technical report.
- reachy-motion-generator-api: a Python client that turns a text prompt into a playable motion.
- reachy-animation: the animation engine that blends generated gestures, idle breathing, and speech-driven head movement.
- reachy-explain: the presentation harness and agent skill behind the launch video. Ask a coding agent to build a talk, and Reachy delivers it with a voice, generated motions, and synchronized slides.
The models are on Hugging Face: 0.8B for speed, 4B for a balance of speed and quality, and 27B for the best motion quality.
I’m also releasing the Reachy Mini Massive Motion Library: 10,872 generated motion episodes, roughly 15 hours of trajectories with rendered videos, packaged as a LeRobot dataset.
If you just want to see the robot move, try the browser demo.

This article is less of a technical report and more of a storytelling article on how I solved problems and came up with ideas. If you want to just read the technicals, feel free to read it here.
How I ended up building an expressive action model
This started on a random weekend. I wasn’t trying to benchmark Astra. I wanted to train a different kind of action model.
In January 2025, Apple published ELEGNT, a paper about a lamp-like robot and how its movements could serve both functional and expressive purposes. What stuck with me was the distinction between completing a task and communicating something through the way you complete it.

Think about closing a door. You can close it quietly because someone is sleeping, or slam it because you want everyone in the room to know you’re angry. The door ends up closed either way. The trajectory carries the difference.
A lot of the action-model work I follow focuses on whether the robot successfully does the task. I’m interested in what happens when we also care about how the movement reads to the person standing next to it.
Then a funny idea popped into my head: what if I trained an end-to-end emotes model?
Give it “heartbroken,” “a cat stalking prey,” or “a sleepy toddler fighting to stay awake,” and have it produce the motion directly.
Two questions immediately got in the way:
- where do I get the data
- and what do I train?
Bootstrapping data
Expressive motion data is hard to find, especially for a robot. Reachy Mini had a somewhat useful starting point: Pollen’s library of 85 recorded emotions, plus 19 dances (about nine minutes of motion).
Nine minutes is very little data, so I was finding an adjacent pre-trained model that could transfer to expressive motion without much fine-tuning. My first thought actually wasn’t an LLM. Inspired by my recent work on action chunks modifications, I tried a music-generation backbone.
My reasoning was that music models learn temporal structure: rhythm, build-up, release, repetition. Maybe some of that would transfer to expressive motion.
Through some searching, I found and fine-tuned ACE-Step’s 2.4B music generation model, directly tuning its audio channel outputs. Naively, if you think about it, the model output audio channels based on text, it should do this task pretty well. In reality, it could learn motion patterns, but couldn't get those patterns to follow my text conditioning. For example, in one check, when told to raise the antennas in a happy motion, none of 144 generated samples raised the antennas.
Fine-tuning a music generation backbone instead of a VLM would have deserved its own article if it hadn’t been absolutely awful lol. Maybe someone smarter than me can make this work.

At first, I blamed the data. With so few clips, memorization was almost inevitable. If I could produce enough good examples, perhaps I could train my way out of it. So I tried asking Astra to choreograph motions to synthesize more data.
The first outputs were so impressive that I stopped and reconsidered the whole approach. Suddenly I could ask for situations that were nowhere in the recorded library and get something recognizably related to the prompt. This was zero-shot, the LLM was extremely good at spatial reasoning.

Fueled by this revelation, I went full throttle and expanded the experiment to 5,000 scenarios, eventually producing 10,000 Astra-authored recipe samples: each scenarios have 2 versions, a dull one and a lively one.

Generating high-frequency trajectories
Making the robot move with Astra is not as simple as telling Astra to generate 25Hz motions. This is:
- Wasteful: I'm not token rich.
- Time-consuming: The control rate of the robot is 25 Hz, meaning 25 full robot states every second. A motion that is 5-second long would consume 125 full robot states, costing roughly 1000 tokens. If the model outputs at 100 token/s, the time to generate that would be 10s. Real-time is off the table.
Basically, asking an LLM to write all of those numbers directly is expensive, slow, and an awkward way to describe a gesture, so I task the LLM to only write a rough motion plan, which is then refined to 25 Hz by a flow matching transformer head.

This is like how an animator animates with key frames. I gave the LLM a similarly compact language. It has three commands: go, hold, and osc. Move toward a pose, hold it, or oscillate a channel.

A recipe can say: tilt back slowly, raise the antennas, pause with tension, snap forward, then settle. It specifies head angles, height, antenna positions, body rotation, and an “energy” value controlling how much fast detail should accompany the posture. That turns a large numerical sequence into a short piece of choreography. It also makes the output easy to validate before using it.
The recipe expands into sparse keyframes. For the served system, those keyframes are spaced every quarter-second, then interpolated to the motion’s 25 Hz frame rate. But interpolation alone doesn’t give me everything I want. It gives me the broad motion. The little overshoots, antenna flicks, and variations in timing are part of what makes the robot look alive.
That is where the second model comes in. I trained a small flow-matching transformer to generate full motion conditioned on the plan. It starts from noise and produces the nine-channel trajectory, using the interpolated keyframes as guidance.


The useful trick is how I got the second model's training pairs. I took each real Pollen animation and extracted a simplified plan: its slower posture changes, plus a measure of the faster detail left over. Furthermore, each recording from Pollen can be described differently, so even though I have 104 original recordings, I can vary the descriptions of them to more training samples.
The generator learns to go from the simplified description back to detailed movement. It never has to see text. It doesn’t need to learn what “heartbroken” means from nine minutes of recordings. The language model handles that part.
At this point, the project had become a two-stage system: a language model choreographs an expression, and a small motion model performs it in Reachy’s style.
Evals
To improve a model, you need a way to measure whether it's good or not. In fact, none of this could have been done if I didn't devise a way to measure the performance so I can let Claude auto-research model improvements.
I started with numerical comparisons: posture, antenna droop, head height, movement speed, and similarity to recorded clips. Those checks were useful. Given a plan extracted from a held-out real motion, the generator’s output matched the correct reference more closely than the other 11 held-out clips about 89% of the time.
However, this only tells me the generator can use a good plan, it doesn’t tell me that a language model’s interpretation of an emotion is the right one. And aggregate metrics can miss something extremely obvious.
For example, my early local planners sneezed upward. They had the general posture and energy of a sudden event. Their losses and average descriptor scores looked reasonable. But instead of tilting back and snapping forward-down, they threw the head up.


The model has produced a perfectly plausible motion for the wrong event.
To solve this problem, I added concrete behavioral probes: does the sneeze release move the head down? Does “nodding yes” actually oscillate in pitch? Does the sleepy toddler droop and then recover? In total, there are 16 probes split between concepts excluded from training and skills with related training examples.

These checks gave me a much more useful signal than an average agreement score. But they only check specific behaviors in the expanded recipes, including eight probes whose concepts were excluded from training. They don’t establish broad out-of-distribution performance or measure how natural and expressive the final generated motion looks. For that, I still watch the robot.
Teaching Astra better taste
Fueled by this revelation, I went full throttle and expanded the experiment to 5,000 scenarios, eventually producing 10,000 Astra-authored recipe samples: each scenarios have 2 versions, a dull one and a lively one.
If you recall this line from before, you might ask: why the two versions?
When Astra first generated the recipes, they were precise, but many had long, frozen holds between poses. They described the action correctly while looking a little too much like a sequence of instructions.
So I hand-picked a subset of motions I really like (roughly 500 of them) and told Astra to improve other samples based on them. That is why I have 10,000 Astra-authored recipe samples across 5,000 scenarios.
That rewrite made a noticeable difference in the distribution of the motion compared.

Distillation and other end-to-end attempts
Meanwhile, I hadn’t completely abandoned the original end-to-end idea.
I tried models trained directly on real motion, models trained on generated motion, and a VLA-style setup with a language backbone and a flow-matching action expert, as in π0.

Some improved on the numerical metrics. The distilled end-to-end models could reproduce familiar prompts reasonably well. But on unusual prompts, they tended to lose the timing, amplitude, or structure that made the planner’s choreography interesting. The VLA-style model was the exception: it kept the planner’s expressive amplitude on unusual prompts, but its motion was still judged worse than the planner and generator.



For this project, the planner-plus-generator system remained the one I preferred to watch. It's much more expressive as you can see above.
So the next question became: could I make the planner small and fast enough to use interactively? That is how the Astra-distillation detour happened.
Instead of asking a small model to learn language and robot motion together, I fine-tuned an existing language model to write recipes. Its training target is a short JSON answer containing one sentence of intent and the choreography. The released planners come in three sizes: 0.8B, 4B, and 27B. The smallest model is useful for quick, single-feeling reactions. The larger models handle sequences of events better: anticipation, release, recovery, or fighting sleep and waking yourself back up. They share the same motion generator.


Using the previously mentioned evals, the 27B reached 96% on the eight excluded-concept probes and 97% on the skill probes. It got the sneeze direction right in all 24 samples across the two sneeze prompts.
I wouldn’t read that as “96% of emotions solved.” It means the model passed those particular checks, and the suite is already close to its ceiling.
Making the model real-time
My speed target is around 100 ms. I wanted a robot to react during a conversation without an awkward pause.
Originally, a full motion on average costs 2.96s to generate.

To drop the latency, I used FP8 weights and speculative decoding, then fine-tuned the model’s multi-token-prediction head on its own recipe outputs so it could draft them more effectively. The motion generator also turned out to need very few sampling steps. Reducing it to eight brought its part of the work down to around 20 ms.
The complete server pipeline took approximately 170 ms with the 0.8B planner, 290 ms with the 4B, and 790 ms with the 27B, on an RTX PRO 6000.
I didn’t quite hit the original 100 ms target, but the smaller models are fast enough to make the interaction interesting. Besides, 200ms is our perception of real-time.
Keeping Reachy alive between gestures
Generating a good clip still only solves part of the experience. A robot that performs a beautiful gesture and then becomes completely motionless looks strange. So does one that snaps from the end of one clip to the beginning of another.
I built reachy-animation to handle the motion between motions. It crossfades gestures into idle breathing and adds small head movements driven by the speech audio. While a new gesture is being generated, the robot keeps moving.



The generated clips are 25 Hz; the animation engine samples and blends them into a continuous stream of poses, 60 times a second by default.
The API library makes the two pieces easy to connect. Give it a prompt such as “proud. You finally solved the puzzle,” and it returns a clip the animator can play. Requests can run in the background, so waiting for the model doesn’t stop the animation loop.
You can try the library here.
The best demo ever
For the launch, I wanted Reachy to explain all of this itself. That became reachy-explain: a presentation harness with an agent skill that guides a coding agent through writing the script, generating the voice and motions, and building the slides.
In the short launch talk, Reachy explains the planner and generator, performs a sneeze, and attempts a more involved sequence: follow an imaginary fly, recoil when it lands on its nose, then regain its composure.

I find what I've created to be the best expressive robot demo I've seen. You can try or adapt this here.
Trying this out
If you want to try it, the browser demo is the easiest place to start. You can watch the prepared examples or connect it to a running generator server. If you have a Reachy Mini, the API client and animation engine provide the path from a prompt to playback.

I’ve also released the teacher recipes and their generated trajectories, with MuJoCo videos, as the Reachy Mini Massive Motion Library: roughly 15 hours of synthetic motion. Those are generated examples, not 15 hours of new robot recordings.
There is plenty left to improve. The generator has a tiny amount of real motion to learn from. The planner inherits its teachers’ taste. Long stories are harder than short reactions, and generating consecutive clips that naturally continue one another is still an open problem here.
I think the biggest opportunity is adapting this pipeline to other robots: starting with a small library of recorded movements and expanding it into a much wider range of expressive behaviors.
Citation
If you use this project in your work, you can cite it as:
@misc{pham2026reachymotion,
author = {Pham, Binh},
title = {Reachy Mini Text-to-Motion},
year = {2026},
url = {https://github.com/pham-tuan-binh/reachy-motion-generator},
note = {Code, models, and technical report}
}