EG
POT-VLA: Persistent 3D Object Tokens Boost Humanoid Loco-Manipulation Success from 39 to 71/80
ResearchJuly 22, 2026Embodied Global

POT-VLA: Persistent 3D Object Tokens Boost Humanoid Loco-Manipulation Success from 39 to 71/80

A team led by Kai Chen introduces Persistent Object Tokenization (POT), a closed-loop approach that maintains role-indexed 3D object records across movement, contact, and occlusion to anchor both action generation and geometric verification. On Unitree G1, POT-VLA improves a direct GR00T-N1.7 baseline from 39/80 to 71/80 successes.

#POT-VLA#Persistent Object Tokenization#humanoid VLA#Unitree G1#arXiv:2607.18016#loco-manipulation
Reading in English

A Core Challenge in Humanoid VLA

Vision-Language-Action (VLA) policies have emerged as a promising foundation for general-purpose robot control. Yet long-horizon humanoid loco-manipulation tasks expose a fundamental gap: the object state used to condition whole-body action can diverge from the state needed to verify whether the action actually achieved the intended physical relation. As the robot moves, makes contact, and experiences occlusion, its internal representation of task objects drifts — undermining both execution accuracy and success-rate evaluation.

This “object-state divergence” problem is why VLA policies that look impressive in short demos often fail in multi-step, real-world tasks. A team led by Kai Chen proposes a clean architectural fix: Persistent Object Tokenization (POT).

How POT-VLA Works

POT maintains role-indexed 3D object records built from RGB-D observations and continuously updated as the robot acts. These persistent records serve two purposes simultaneously:

  1. Action generation — the 3D object tokens are injected into a whole-body action expert as grounded state conditioning.
  2. Geometric verification — the same records support predicate-based success checks (e.g. “is the cup on the table?”).

The result is a closed-loop execution system in which object state is both actionable and verifiable. When the object record diverges from reality after an action, the verifier catches the failure and triggers recovery — instead of silently compounding errors across a long task sequence.

Results on Unitree G1

The authors instantiate the approach as POT-VLA and evaluate it on a Unitree G1 humanoid across eight real-world task families, using a matched GR00T-N1.7 baseline as the comparison point.

  • Baseline GR00T-N1.7: 39 / 80 successes
  • POT-VLA: 71 / 80 successes — an 82% relative improvement

On an external Being-0-aligned reference set of service tasks, POT-VLA achieves 44 / 50 successes versus the 37 / 50 reported in the Being-0 paper.

The largest gains appear precisely on tasks requiring maintained 3D spatial relations — pick-and-place with precision constraints, opening doors while holding objects, and multi-step rearrangement tasks. This confirms that persistent, object-centered state is the missing abstraction for verifiable humanoid VLA execution.

Why It Matters

POT-VLA matters for three reasons:

  1. Architectural simplicity. Rather than building a separate verifier, the same 3D object tokens drive both acting and checking — reducing engineering complexity.
  2. Real-robot validation. Results are on physical Unitree G1 hardware, not just simulation, with eight diverse task families.
  3. Measurable, reproducible gains. Going from 39/80 to 71/80 is a large, quantifiable leap that other VLA systems can directly benchmark against.

The paper (arXiv:2607.18016) was submitted on July 20, 2026 by Peng Ren, Haoyang Ge, Jiang Zhao, Cong Huang, Yukun Shi, Pei Chi, and Kai Chen. It joins a fast-growing body of work treating “object persistence” as the next critical capability gap in embodied AI — and suggests a simple, token-level fix may go a long way.

Language: English- Showing content in English