PAPER READOUT / READING LIBRARY
論文閱讀庫
共 2 篇 · 雙語摘錄、閱讀批註與圖表導讀
-
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
FlashAttention-3 利用 Hopper GPU 的非同步資料搬移與矩陣運算,並處理 FP8 的資料布局和量化誤差,使精確注意力運算更快。
-
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
FlashAttention-2 在維持精確注意力計算與線性額外記憶體需求的前提下,重新安排演算法、thread block 與 warp 的工作,提升 GPU 吞吐量。