[js/webgpu] support FlashAttention-2 for attention operator #22915

xhcao · 2024-11-21T08:30:46Z

Description

Motivation and Context

xhcao · 2024-11-21T08:43:07Z

The FlashAttention-2 algorithm is based on the paper https://tridao.me/publications/flash2/flash2.pdf.
Memory increase quadratically in the sequence length in original algorithm, and need a lot of global memory read and write accesses. If sequence is large, for example running stable diffusion 2.1 with attention nodes, original algorithm will lead chrome crash for out of memory and inefficient.
Please take a look, @axinging @hujiajie @jiechen0826

guschmue · 2024-11-21T16:54:02Z

/azp run ONNX Runtime Web CI Pipeline,Windows GPU CI Pipeline,Linux Android Emulator QNN CI Pipeline

guschmue · 2024-11-21T16:54:11Z

/azp run Linux CPU CI Pipeline,Linux CPU Minimal Build E2E CI Pipeline,Linux GPU CI Pipeline,Linux GPU TensorRT CI Pipeline,Linux OpenVINO CI Pipeline,Linux QNN CI Pipeline,MacOS CI Pipeline,Windows ARM64 QNN CI Pipeline,Windows CPU CI Pipeline

azure-pipelines · 2024-11-21T16:54:16Z

Azure Pipelines successfully started running 1 pipeline(s).

guschmue · 2024-11-21T16:54:18Z

/azp run Windows GPU TensorRT CI Pipeline,onnxruntime-binary-size-checks-ci-pipeline,orttraining-linux-ci-pipeline,orttraining-linux-gpu-ci-pipeline,orttraining-ortmodule-distributed,Windows x64 QNN CI Pipeline,Big Models

azure-pipelines · 2024-11-21T16:54:23Z

Azure Pipelines could not run because the pipeline triggers exclude this branch/path.

guschmue · 2024-11-21T16:54:24Z

/azp run Windows GPU CUDA CI Pipeline,Windows GPU DML CI Pipeline,Windows GPU Doc Gen CI Pipeline

azure-pipelines · 2024-11-21T16:54:32Z

Azure Pipelines could not run because the pipeline triggers exclude this branch/path.

azure-pipelines · 2024-11-21T16:54:32Z

Azure Pipelines successfully started running 1 pipeline(s).

xhcao · 2024-11-22T04:14:12Z

Sorry for the error, I wanted to keep the names as the paper, but the variables‘ names mismatched rules. I had already modified the names. Thanks.

axinging · 2024-11-25T05:57:22Z

js/web/lib/wasm/jsep/webgpu/ops/attention.ts

+      Q_i[local_id.y][u32(${workgroupSize[0]} * tile) + local_id.x] = Q[offset + local_id.y * uniforms.d + u32(${workgroupSize[0]} * tile) + local_id.x];
+    }
+
+    for (var j = 0; j < ${tC}; j++) {


Turn tC to uniform or add it to hint?

Thanks. I also noticed this issue. I added it to uniform.

axinging · 2024-11-25T05:58:30Z

js/web/lib/wasm/jsep/webgpu/ops/attention.ts

+    context.inputs[4] === undefined &&
+    context.inputs[5] === undefined
+  ) {
+    return applyFlashAttentionV2(context, q, k, v, params, attributes);


Any existing case cover this branch?

Currently, I also tested it on sd2.1. From the conditions, we know that the input data size is large, so I am not sure it is reasonable to add an unit test here.

jchen10 · 2024-11-25T06:33:25Z

The FlashAttention-2 algorithm is based on the paper https://tridao.me/publications/flash2/flash2.pdf. Memory increase quadratically in the sequence length in original algorithm, and need a lot of global memory read and write accesses. If sequence is large, for example running stable diffusion 2.1 with attention nodes, original algorithm will lead chrome crash for out of memory and inefficient. Please take a look, @axinging @hujiajie @jiechen0826

@xhcao my github name is jchen10.

xhcao force-pushed the flash-attention-2 branch from 1eb8d05 to 1aaa47e Compare November 21, 2024 08:34

guschmue added the ep:WebGPU ort-web webgpu provider label Nov 21, 2024

[js/webgpu] support FlashAttention-2 for attention operator

c936365

xhcao force-pushed the flash-attention-2 branch from 1aaa47e to c936365 Compare November 22, 2024 04:10

axinging reviewed Nov 25, 2024

View reviewed changes

Add tc to uniform and resolve hard numbers

735ccf2

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[js/webgpu] support FlashAttention-2 for attention operator #22915

[js/webgpu] support FlashAttention-2 for attention operator #22915

xhcao commented Nov 21, 2024

xhcao commented Nov 21, 2024

guschmue commented Nov 21, 2024

guschmue commented Nov 21, 2024

azure-pipelines bot commented Nov 21, 2024

guschmue commented Nov 21, 2024

azure-pipelines bot commented Nov 21, 2024

guschmue commented Nov 21, 2024

azure-pipelines bot commented Nov 21, 2024

azure-pipelines bot commented Nov 21, 2024

xhcao commented Nov 22, 2024

axinging Nov 25, 2024

xhcao Nov 25, 2024

axinging Nov 25, 2024

xhcao Nov 25, 2024

jchen10 commented Nov 25, 2024

[js/webgpu] support FlashAttention-2 for attention operator #22915

Are you sure you want to change the base?

[js/webgpu] support FlashAttention-2 for attention operator #22915

Conversation

xhcao commented Nov 21, 2024

Description

Motivation and Context

xhcao commented Nov 21, 2024

guschmue commented Nov 21, 2024

guschmue commented Nov 21, 2024

azure-pipelines bot commented Nov 21, 2024

guschmue commented Nov 21, 2024

azure-pipelines bot commented Nov 21, 2024

guschmue commented Nov 21, 2024

azure-pipelines bot commented Nov 21, 2024

azure-pipelines bot commented Nov 21, 2024

xhcao commented Nov 22, 2024

axinging Nov 25, 2024

Choose a reason for hiding this comment

xhcao Nov 25, 2024

Choose a reason for hiding this comment

axinging Nov 25, 2024

Choose a reason for hiding this comment

xhcao Nov 25, 2024

Choose a reason for hiding this comment

jchen10 commented Nov 25, 2024