Figure AI, a robotics firm, has recently showcased two F.03 humanoid robots autonomously resetting an entire bedroom, including collaboratively making a bed, using only visual cues to synchronize their movements. This demonstration signifies a groundbreaking achievement as it is the first time a single neural network has been trained to enable collaborative tasks across multiple humanoid robots solely based on visual input guiding their actions.
The robots are driven by Helix-02, a unified Vision-Language-Action system governing the complete functionalities of each robot. Unlike conventional robotic systems that rely on separate planners, message exchanges, or a central controller, these two humanoid robots perceive their surroundings through their individual cameras and interpret each other’s intentions solely through motion cues. Simple gestures like a nod, a change in posture, or the position of an arm enable them to stay coordinated.
In the demonstration, the robots are seen performing various tasks such as opening doors, hanging clothes on a coat stand, tidying up objects, and collaborating to make the bed. The bed-making task is particularly challenging, involving lifting, unfolding, spreading, folding, and smoothing a duvet with precision, ensuring wrinkles are corrected and edges are neatly aligned. Notably, all actions occur in real-time without any remote control or human intervention.
The underlying Helix-02 system, although not initially designed for bedroom tasks, is a versatile learned policy that enhances its capabilities with more data input. Previously, this system facilitated a Figure robot in loading a dishwasher in a kitchen within four minutes and enabled an F.03 robot to tidy a living room by performing various chores. The bedroom reset demonstrates another layer of capability achieved without altering the core algorithm.
The complexity of the bed-making sequence arises from three main challenges. Firstly, coordinating two robots in the same space presents interdependent tasks where each robot’s actions influence the other’s tasks. Secondly, the duvet, being a flexible object without defined boundaries, requires the robots to predict each other’s movements while maintaining physical contact points. Lastly, the sequence demands swift execution, necessitating seamless transitions between different manipulation techniques without pauses or scripted handovers.
Brett Adcock, CEO of Figure AI, highlighted that the robots coordinated their actions solely through visual cues like head nods, emphasizing the fully autonomous nature of the task conducted at normal speed without any external control.
The company views this demonstration as a significant stride towards a future where intelligent humanoid robots can collaborate efficiently in various environments such as homes, warehouses, and factories. They envision these robots handling shared objectives in dynamic spaces with moving entities. As the robots continue to learn and evolve with more data, Figure AI anticipates expanding their capabilities and is actively seeking new talent to join the team.
