Developed by the University of Hong Kong’s MMLab alongside nearly 20 international academic institutions, RoboDojo pushes robots beyond simple demonstration-based training. VPP2 distinguishes itself from standard video generation models by focusing on object dynamics and the translation of high-level instructions into precise physical movements. By integrating video prediction with action generation, the system allows robots to navigate unfamiliar surroundings and complex manipulation tasks with increased reliability.
Performance metrics highlight the model’s versatility. On the ALOHA platform, VPP2 outperformed existing baselines in nine out of ten task categories, achieving a 58.5% success rate. The integration of a vision-language model for high-level task planning further improved performance, more than doubling success rates on the LIBERO-Pro benchmark from 27.6% to 57.6%.
This success marks the fourth time in 2026 that ROBOTERA has claimed a benchmark championship in embodied intelligence. Beyond academic testing, the firm is currently integrating these advancements into logistics operations, with humanoid units deployed at more than 10 facilities for China Post and SF Express. The company has since made the VPP2 model architecture available to the public via GitHub.



Comments (0)
No comments yet. Be the first!