Researchers at UCLA's Robotics and Mechanisms Laboratory (RoMeLa) have published RAVEN, a hierarchical reinforcement learning-MPC framework designed to make humanoid robot navigation more robust in dynamic, real-world environments. The paper, led by Ruochen Hou with Dennis W. Hong as senior author, appeared on arXiv on July 17, 2026.
Humanoid navigation requires long-horizon planning while respecting short-horizon dynamic and safety constraints. Classical visibility-graph planners combined with model predictive control (MPC) can efficiently generate collision-free trajectories, but their performance depends on manually tuned parameters and accurate system modeling. On real robots, control delays, state-estimation noise, and locomotion uncertainties can cause overshoot and constraint violations even when the nominal path is geometrically optimal.
Instead of using learning to tune cost weights — or replacing planning entirely with an end-to-end policy — RAVEN employs reinforcement learning to adapt the geometric construction of the visibility-graph planner itself, modifying obstacle inflation and related graph parameters. By directly reshaping free-space geometry, the learned planner alters the topology of the global path to compensate for delay and tracking imperfections. A collision-free MPC layer then tracks the planned trajectory while explicitly enforcing velocity bounds and obstacle-avoidance constraints.
Trained under realistic delays and observation noise, RAVEN was evaluated against a manually tuned visibility-graph MPC baseline and a pure RL navigation policy. The results show reduced overshoot near obstacles, improved robustness in narrow passages, and more reliable navigation under delay and noise.
The work points to an interpretable middle ground between hand-engineered navigation stacks and end-to-end learning — a relevant direction as humanoid developers push robots from controlled demos into cluttered warehouses and factory floors where sensing and actuation are never perfect.


