原文 SOURCE
More fundamentally, FlashAttention-2’s algorithm adheres to a simplified synchronous model and makes no explicit use of asynchrony and low-precision in its design. Asynchrony is a result of hardware specialization to accelerate the most important operations in a ML workload: specific hardware units performing matrix multiplication (Tensor Cores) or memory loading (Tensor Memory Accelerator – TMA), separate from the rest of the CUDA cores performing logic, integer, and floating point computation. Low precision such as FP8 in Hopper and FP4 in Blackwell, continuing the trend of FP16 (Pascal in 2017) and BF16 (Ampere in 2020), is a proven technique to get double or quadruple throughput for the same power and chip area. We review the capabilities afforded by Hopper in these directions in § 2.2. The technical challenge is to redesign FlashAttention-2 to make use of these hardware features: asynchrony requires overlapping computation between matmul and softmax even though one depends on the output of the other, and low-precision requires care to minimize quantization error, especially in the case of outlier features in LLMs [20, 54].
繁體中文 TRANSLATION
更根本的是,FlashAttention-2 的演算法遵循簡化的同步模型,在設計上並未明確使用非同步能力與低精度。非同步能力來自硬體的專門分工,以加速機器學習工作負載中最重要的操作:由特定硬體單元執行矩陣乘法(Tensor Cores)或記憶體載入(Tensor Memory Accelerator,TMA),有別於執行邏輯、整數及浮點計算的其他 CUDA cores。Hopper 的 FP8 與 Blackwell 的 FP4 等低精度格式,延續 FP16(2017 年的 Pascal)及 BF16(2020 年的 Ampere)的趨勢,是在相同功耗和晶片面積下取得兩倍或四倍吞吐量的已驗證技術。我們在 § 2.2 回顧 Hopper 在這兩方面提供的能力。技術上的挑戰是重新設計 FlashAttention-2 以使用這些硬體特性:非同步執行要求矩陣乘法與 softmax 重疊,儘管後者依賴前者的輸出;低精度則必須謹慎控制量化誤差,尤其是在大型語言模型具有離群特徵的情況下 [20, 54]。