首页|PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind

PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind

来源：

英文摘要

Multi-agents has exhibited significant intelligence in real-word simulations with Large language models (LLMs) due to the capabilities of social cognition and knowledge retrieval. However, existing research on agents equipped with effective cognition chains including reasoning, planning, decision-making and reflecting remains limited, especially in the dynamically interactive scenarios. In addition, unlike human, prompt-based responses face challenges in psychological state perception and empirical calibration during uncertain gaming process, which can inevitably lead to cognition bias. In light of above, we introduce PolicyEvol-Agent, a comprehensive LLM-empowered framework characterized by systematically acquiring intentions of others and adaptively optimizing irrational strategies for continual enhancement. Specifically, PolicyEvol-Agent first obtains reflective expertise patterns and then integrates a range of cognitive operations with Theory of Mind alongside internal and external perspectives. Simulation results, outperforming RL-based models and agent-based methods, demonstrate the superiority of PolicyEvol-Agent for final gaming victory. Moreover, the policy evolution mechanism reveals the effectiveness of dynamic guideline adjustments in both automatic and human evaluation.

作者：Yajie Yu、Yue Feng

作者单位：

学科分类：计算技术、计算机技术

推荐引用：Yajie Yu,Yue Feng.PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind[EB/OL].(2025-04-20)[2025-05-25].https://arxiv.org/abs/2504.15313.点此复制

PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind

PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind

评论