Barcelona during spring has a way of making everything feel alive, sunlight bouncing off glass buildings, trees starting to bloom, and, in our case, humanoid robots taking their first (slightly wobbly) steps.
My name is Josep, a Doctoral Candidate in Emerging Digital Technologies at Scuola Superiore Sant’Anna in Pisa, and I found myself in the middle of all this when I recently had the opportunity to spend three days attending The Construct’s course on Humanoid Robot Reinforcement Learning, with the support of PAL Robotics. It turned out to be one of those uncommon experiences where learning, experimentation, and fun all come together.
The promise of the bootcamp was ambitious: go from nothing, to understanding the basics of the reinforcement learning (RL) pipeline in just three days.
Surprisingly, it was delivered.
Day 1: Learning Humanoid Locomotion
Using Isaac Lab, we trained reinforcement learning policies that allowed the humanoid to walk in simulation. Defining our own reward functions and watching the policy improve over time, from unstable, almost random movements to something that resembled balance and coordination, was incredibly satisfying.
By the end of the day, we had something that worked in simulation. Not perfectly but knowing that half of us didn’t even know what a cost function was when we started, it felt like a real success.
Day 2: From Locomotion to Whole-Body Skills
On day two, instead of training the robot by explicitly defining cost and reward functions, we trained it using motion recordings.
We were provided with a bank of pre-recorded motions that had been segmented and translated into joint trajectories, making them usable as training data for humanoid robots.
Each of us selected a short fragment, whether a dance move, a martial arts kick, or a particularly challenging jump, and trained a model to reproduce it.
After training our individual models, we compared results, discussed what worked and what didn’t, and selected some of the most successful (and entertaining) ones. These were then deployed on the real robot.
Seeing those movements executed physically, after only a few hours of work, was easily one of the highlights of the course.

Day 3: Towards Generalization and Research
The final day shifted gears toward more research-oriented ideas.
We explored Vision-Language-Action models and experimented with MuJoCo-based environments. This wasn’t just about making a robot perform a task; it was about understanding how these systems might learn to operate under ambiguous instructions: “Move the box to the left” (how much to the left?), “bring me a glass of water” (how full? which glass?), and so on.
These questions highlighted how far current systems have come, and how much remains to be solved.

Delivery of the certificate
Conclusion
In just three days, we went from basic concepts to training and deploying behaviors on real humanoid robots. More importantly, we experienced the full process, what works, what fails, and how to iterate.
For me personally, the connections to my own research were impossible to ignore. My work focuses on developing adaptive control strategies for exoskeletons to better assist users with varying needs and impairments, and the tools and frameworks we explored here are ones I fully intend to bring back to the lab. It’s a reminder that the gap between theory and real deployment can close faster than you’d expect — when you have the right environment to experiment in.
If you’re curious about my research or want to connect, feel free to look me up through the AERIALIST Doctoral Network.