[论文] Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computat…

## 论文概要 **研究领域**: ML **作者**: Yujie Zhang, Huiying Lan, ...

论文概要

研究领域: ML 作者: Yujie Zhang, Huiying Lan, Ehsan Aghapour 发布时间: 2026-09-06 arXiv: 2509.04277

中文摘要

随着边缘深度学习应用变得日益复杂,在异构片上系统(SoC)上优化性能面临独特挑战。传统流水线技术将计算分布在不同片上处理单元之间,虽对吞吐量有效,但无法满足具有复杂依赖关系和大量算子并行性的现代神经网络的延迟需求。利用算子并行性在多个处理单元上实现并发执行有降低推理延迟的潜力。然而,优先考虑流水线或并行执行通常需要权衡,优化一个性能指标会对另一个产生不利影响。本文介绍了Para-Pipe,一种分层映射框架,将流水线架构内的阶段内和阶段间算子并行性集成在一起。Para-Pipe通过在流水线阶段内和跨阶段选择性微调并行级别来权衡吞吐量和延迟。该策略可显著降低处理器间通信开销,大幅提高能效。我们的评估表明,Para-Pipe在配备ARM big.LITTLE CPU和GPU的Amlogic SoC,以及具有深度学习加速器和两个DSP的黑芝麻科技SoC上,生成了多个帕累托最优配置,实现了吞吐量与延迟的平衡。更重要的是,Para-Pipe在Amlogic SoC上的吞吐量优化配置相比纯流水线策略平均能效提升11.0%,相比非流水线并行执行提升23.3%。

原文摘要

As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential in leveraging operator parallelism to enable concurrent execution across multiple processing units, thereby reducing inference latency. However, prioritizing pipelining or parallel execution often necessitates a compromise, where optimizing one performance metric adversely impacts the other. This paper introduces Para-Pipe, a hierarchical mapping frame…

自动采集于 2026-09-07

#论文 #arXiv #ML #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1