SheepNav
新上线10天前0 投票

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents

arXiv:2605.20246v1 Announce Type: new Abstract: Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple turns of visual perception and action execution. However, existing methods still rely primarily on Supervised Fine-Tuning (SFT) with expert demonstrations, while the advanced reinforcement learning (RL) algorithm, specifically Group Relative Policy Optimization (GRPO), has not been effectively empl

延伸阅读

  1. 戴尔新款XPS 13售价599美元,挑战MacBook Neo,保留高端特性
  2. 戴尔 XPS 13 (2026) vs. MacBook Neo:两款平价笔记本对比,我选这款
  3. 艾琳·布罗克维奇瞄准数据中心保密问题
查看原文