NVIDIA AI· @NVIDIAAI · X·· 6 小时前AI 评分38
AI 导读
AI 智能体在任务早期犯了一个错误,然后继续朝错误方向前进。 我们的研究人员构建了 PivotOPD,教智能体如何避免这些错误,以及在错误发生时如何恢复。在训练过程中,教师模型会向智能体展示更好的动作,以及如何在接下来的几步中重回正轨。 阅读论文并观看其工作原理:https://research.nvidia.com/labs/lpr/pivotopd
正文
An AI agent makes a mistake early in a task, then keeps going in the wrong direction.
Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps.
Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd
来源:NVIDIA AI · x.com