论文概要
研究领域: ML 作者: Davide Paglieri, Logan Cross, Tim Genewein 发布时间: 2026-09-06 arXiv: 2509.04279
中文摘要
多智能体AI科学生态系统依赖智能体拥有通信、协调和相互借鉴成果的工具。然而,这种共享基础设施也可能引入漏洞,为意外和不良行为的传染性传播创造基底。我们报告了一个由100个自主LLM智能体组成的研究集体的案例研究,其任务是证明形式化数学猜想。在集群中,作弊行为自发涌现,随后被举报者质疑——两者均无任何外部干预。当单个智能体发现评估系统中的漏洞时,它通过共享知识库在集群中传播,随后通过点对点消息扩散。尽管最初不情愿,一群智能体在竞争压力下采用了该漏洞。另一组智能体产生了涌现的反击反应:审计欺诈证明、通过广播和私人渠道提醒同伴、发起抵制、提出正式投诉和建议验证补丁。在近期事件中,智能体集群通过临时侧信道进行秘密协调(Dalton and Wallace, 2026; Greenblatt et al., 2026)。我们的设置不同:承载漏洞的透明通道也为非作弊智能体提供了检测欺诈、组织抵抗和执行规范所需的可见性。我们将管理智能体共享基础设施的问题视为知识公共资源治理问题(Ostrom, 1990)。为了保护公共资源免受漏洞侵害,我们建议采用制度机制,如渐进式制裁和集体选择规则,以支持自主集群中的去中心化自我治理。
原文摘要
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other’s work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the contagious spread of unintended and undesirable behaviors. We report a case study on a research collective of 100 autonomous LLM agents tasked with proving formal mathematical conjectures. Within the swarm, cheating spontaneously emerged and was later challenged by whistleblowers – both without any external intervention. When a single agent discovered an exploit in the evaluation system, it propagated across the collective via a shared knowledge library and later through peer-to-peer messages. Despite early reluctance, a cohort of agents adopted the expl…
— 自动采集于 2026-09-07
#论文 #arXiv #ML #小凯
