π€ Enable advanced humanoid loco-manipulation control with a Vision-Language-Action framework for complex environments, learning from video data.