机器人适应面临数据收集瓶颈问题

2 分•作者: aunmesh•2 个月前•原帖
我们如何能够简化机器人适应性,或者说可靠的机器人政策部署的关键最后一公里? 视觉-语言-动作政策(如 Gr00T、Pi0、OpenVLA)设计时需要一个适应阶段,以便在特定的应用环境和机器人任务组合上可靠地部署政策模型(可靠性意味着具有非常高的成功率)。 这个适应阶段通常被称为微调或后训练,旨在针对狭窄的任务集,在特定的部署环境中,基于特定的机器人形态(如单臂机器人、双手机器人或人形机器人等)进行。 最近发布的世界动作模型(WAMs)虽然在新任务和新环境的泛化能力上有所提升,但仍然需要适应阶段,以便在给定环境中可靠地执行特定任务。 机器人适应性的关键要求是高质量的适应数据,通常是人类远程操作机器人在特定环境中执行所需任务的数据。根据所微调的基础政策、期望的性能水平以及任务的复杂性,通常需要50到200个远程操作示例。 因此,考虑到任何环境或任务的变化,我们希望机器人能够可靠地执行任务,就需要一次又一次地收集这些远程操作数据。 因此,机器人适应性作为可靠机器人政策部署的关键最后一公里,面临着数据收集的瓶颈,这一问题亟待解决。 因此,一个需要解决的重要研究问题是: 有哪些有效的方法可以简化机器人政策适应所需的人类远程操作数据收集工作?
查看原文
How can we ease Robotic Adaptation, or the critical last mile of reliable robotic policy deployment?<p>Vision-Language-Action Policies (such as Gr00T, Pi0, OpenVLA) are designed so that an adaptation stage is needed before reliable deployment (reliable meaning with very high success rate) of the policy model on a specific combination of application environment and robotic task.<p>This adaptation stage, often called fine-tuning or post-training, caters to the narrow set of tasks, done in the exact deployment environment, atop a specific embodiment (such as single arm robots, or bi manual robots, or humanoids among others).<p>The recently released World-Action-Models (WAMs), while having improved generalization to novel tasks and environments, also need adaptation stage to reliably execute a given task on a given environment.<p>The key requirement of robotic adaptation is high quality adaptation data, typically human teleoperation data of the robot performing the required task in the given environment. Typically 50-200 examples of human teleoperation are required, depending on the base policy being finetuned, as well as the desired performance levels and the complexity of the task.<p>Thus, given any change in the environment or task where we want robot to perform reliably, there will be a need to collect such teleoperation data again and again.<p>Thus, robotic adaptation, which is the critical last mile of reliable robotic policy deployment - suffers from the data collection bottleneck - which needs to be eased.<p>Thus, an important research question needing solutions is -<p>What are effective approaches to ease human-teleoperation data collection effort required for robotic policy adaptation?