AI研究・論文解説 FlashAttentionとは?AttentionをIO-awareに高速化する仕組み
FlashAttentionとは、Self-Attentionを近似せず、GPUのHBMとSRAM間のIOを減らして高速・省メモリ化する手法です。tiling、online softmax、recomputation、Block-sparse FlashAttention、実験結果を論文ベースで解説します。
AI研究・論文解説
AI研究・論文解説
AI研究・論文解説
AI研究・論文解説
AI技術・仕組み
AI技術・仕組み
AI技術・仕組み
AI研究・論文解説
AI技術・仕組み
AI技術・仕組み