[论文] Proteus: Incremental Memory Activation for Long-Context Sequence Model…

## 论文概要 **研究领域**: ML **作者**: Reza Bayat, Ali Behrouz, V...

论文概要

研究领域: ML 作者: Reza Bayat, Ali Behrouz, Vahab Mirrokni et al. (4 authors) 发布时间: 2026-08-17 arXiv: 2608.16844

中文摘要

基于注意力的序列模型处理长上下文的二次成本促使了越来越多的记忆模型研究,这些模型可将上下文压缩为紧凑状态。然而,大多数现有记忆模型在整个序列中暴露静态记忆。由于早期token没有压缩压力,它们占用太多自由度并污染记忆状态,为后续上下文留下很少容量,并增加存储内容与到达内容之间的干扰。我们研究了一种增量记忆激活的新范式,其中记忆的有效容量随着上下文增长而逐步扩展。施加早期瓶颈迫使模型更有效地压缩历史,而随时间解锁新容量减少干扰并改善对后续上下文的保留。我们在Proteus中实例化这一范式,这是一个可无缝融入广泛神经记忆架构类别的简单机制。我们将Proteus应用于SOTA模型(SWLA、Comba、Titans、Hope-Attention),在标准语言建模和推理以及长上下文检索和理解上观察到一致的改进,增益随上下文长度增加而增长。

原文摘要

The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a compact state. However, most existing memory models expose a static memory throughout the entire sequence. Because early tokens face no compression pressure, they occupy too many degrees of freedom and pollute the memory state, leaving little capacity for later context and increasing interference between what is stored and what arrives next. We study a new paradigm of incremental memory activation, where the effective capacity of memory is progressively expanded as the context grows. Imposing an early bottleneck forces the model to more effectively compress history, while unlocking fresh capacity over time reduces interfere…

自动采集于 2026-08-19

#论文 #arXiv #ML #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1