[论文] Towards Computational Provenance: Carrying Causal-State Evidence in Ge…

## 论文概要 **研究领域**: NLP **作者**: Benjamin Belay **发布时间**: ...

论文概要

研究领域: NLP 作者: Benjamin Belay 发布时间: 2026-08-17 arXiv: 2608.16868

中文摘要

语言模型的输出本身不能提供关于产生它的内部计算的可验证证据。我们研究计算溯源:生成的文本是否能携带关于哪个因果相关内部状态发生的可检测证据。我们在两种受控架构中测试了这一想法的有界形式:模块化前馈神经网络和基于transformer的模型。两种架构都在相同的算术任务上训练,具有通过两个离散中间状态的强制路径,允许不同的内部路径产生相同的答案。我们故意在这些路径之间切换,认证实际使用的状态,并让该验证状态在生成的文本中确定一个微妙的统计模式,该模式随后可被检测。前馈和transformer系统在公开和单独密封的端到端评估中都通过了所有128对匹配测试,检测器恢复了与认证内部状态相关的信号。这些结果为关于经验证的因果相关的内部状态的信息可以在答案不变的情况下保留在生成的文本中提供了受控的概念验证。

原文摘要

A language model’s output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal state occurred. We test a bounded form of this idea in two controlled architectures: a modular feed-forward neural network and a transformer-based model. Both architectures are trained on the same arithmetic task with a mandatory pathway through two discrete intermediate states, allowing different internal paths to produce the same answer. We deliberately switch between these paths, authenticate the state actually used, and let that verified state determine a subtle statistical pattern in the generated text that can later be detected. The feed-forw…

自动采集于 2026-08-19

#论文 #arXiv #NLP #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1