Improve the precision of the FusedAddRMSNormKernel function #587

Abatom · 2024-11-06T06:56:33Z

When sizeof(T) == 2, the sum of the read input and residual (float x) is split into two parts, high and low 16 bits, and saved to input and residual respectively. Later, input and residual are read out and combined to x, with the aim of improving the precision of the subsequent x * rms_rcp operation.

Increase precision from 1e-2 to 1e-3.

Abatom · 2024-11-06T07:02:13Z

def fused_add_rms_norm(x, residual, weight, eps):
    orig_dtype = x.dtype
    x = x.to(torch.float32)
    x = x + residual.to(torch.float32)
    residual = x.to(orig_dtype)

    variance = x.pow(2).mean(dim=-1, keepdim=True)
    x = x * torch.rsqrt(variance + eps)
    x = x.to(orig_dtype) * weight
    return x, residual

If the function is modified as follows, the output result of the fused_add_rms_norm function will be almost the same as that of FusedAddRMSNormKernel, the precision can reach 1e-20.

def fused_add_rms_norm(x, residual, weight, eps):
    orig_dtype = x.dtype
    x = x.to(torch.float32)
    x = x + residual.to(torch.float32)
    residual = x.to(orig_dtype)

    variance = x.pow(2).mean(dim=-1, keepdim=True)
    x = x * torch.rsqrt(variance + eps) * weight.to(orig_dtype)
    return x.to(orig_dtype), residual

yzh119

Nice contribution, thank you @Abatom !
Left some comments for discussion.

include/flashinfer/norm.cuh

zhyncs · 2024-11-06T09:05:06Z

It's better to add the benchmark result for the new one @Abatom

yzh119 · 2024-11-06T09:19:01Z

@zhyncs we haven't set up a standard benchmark for normalization kernels so I think we can leave it for further work.

One interesting feature to have in flashinfer is to add benchmarking class that returns bandwidth and FLOP utilization like proton. Ideally we can port nvbench to python but I don't have a concrete idea about the amount of work.

Abatom · 2024-11-06T11:17:06Z

@yzh119 The shared memory has already been used in place of global memory, and an global memory read has also been reduced.

yzh119

LGTM, I think this PR is ready to be merged.

Brief note (to remind myself what this PR is doing): keep residual in fp32 in shared memory to increase the numerical accuracy of rmsnorm.

gemma-style rmsnorm kernels (introduced in #477 ) are similar to original rmsnorm kernel, and we should use the same kernel for them. This PR cleans up duplicate code and unifies the kernels for gemma-style and original rmsnorm kernels. The precision improvements (#587, #592) are kept in this PR.

🤖 I have created a release *beep* *boop* --- ## [0.2.0](v0.1.6...v0.2.0) (2024-12-17) [Release Blog](https://flashinfer.ai/2024/12/16/flashinfer-v02-release.html). ### Features * add `rotary_dim` argument to rope APIs for partial apply rope ([#599](#599)) ([eb9bc71](eb9bc71)) * add a `use_softmax` field in variant class ([#533](#533)) ([d81af97](d81af97)) * add an option `non_blocking` to plan function ([#622](#622)) ([560af6f](560af6f)) * add gemma_rmsnorm and gemma_fused_add_rmsnorm ([#477](#477)) ([1a6b17e](1a6b17e)) * add group size 3 to GQA decode dispatch ([#558](#558)) ([6227562](6227562)) * add JIT compilation support for FA3 templates ([#672](#672)) ([d4e8d79](d4e8d79)) * allow the cascade kernels to be executed using varying sequence lenghts ([#627](#627)) ([92ac440](92ac440)) * CUDAGraph compatibility of multi-level cascade inference APIs ([#586](#586)) ([2332e8a](2332e8a)) * fix the maximal grid dimension in prefill planning with CUDA graphs ([#639](#639)) ([86ca89a](86ca89a)) * improve the precision of the FusedAddRMSNormKernel function ([#587](#587)) ([c7dc921](c7dc921)) * JIT compilation ([#507](#507)) ([3613a5b](3613a5b)) * modify group-gemm stage number ([#497](#497)) ([52dab1d](52dab1d)) * non-contiguous query with paged kv cache ([#553](#553)) ([89f2c4a](89f2c4a)) * pass a dynamic token count to the cascade kernels ([#635](#635)) ([5fe9f7d](5fe9f7d)) * simplify prefill JIT compilation ([#605](#605)) ([fe4f898](fe4f898)) * specify gemm backend ([#648](#648)) ([0cc1a51](0cc1a51)) * support cached cos/sin in rope APIs ([#585](#585)) ([83e541d](83e541d)) * support huggingface transformer style rope interface ([#568](#568)) ([4f40420](4f40420)) * support sm90 cutlass group gemm ([#509](#509)) ([794bdda](794bdda)) * torch custom_op fix for rope ([#569](#569)) ([3e104bc](3e104bc)) * torch custom_op support: norm ([#552](#552)) ([f6e0010](f6e0010)) * torch.compile and custom_op support ([#554](#554)) ([9bf916f](9bf916f)) * warmup for jit kernel tests ([#629](#629)) ([8f5f349](8f5f349)) ### Bug Fixes * AOT compiler flags on non-sm90 ([#522](#522)) ([0aa4726](0aa4726)) * batch decode kernel redundant store output to gmem ([#505](#505)) ([90e42a7](90e42a7)) * compatible with torch 2.2 ([#478](#478)) ([ac41d1b](ac41d1b)) * #452 ([b53a46f](b53a46f)) * remove redundant load ([#495](#495)) ([2de16b0](2de16b0)) * update bmm fp8 test ([#487](#487)) ([45eac04](45eac04)) ### Performance Improvements * accelerate JIT compilation speed ([#618](#618)) ([eaf73fd](eaf73fd)) * Dense and sparse customizable flashattention-3 template ([#667](#667)) ([51236c9](51236c9)) * fix prefill kernel performance degradation (step 1) ([#602](#602)) ([595cf60](595cf60)) * fix the performance issue of `append_paged_kv_cache` ([#588](#588)) ([e15f7c9](e15f7c9)) * improve parallelism in RoPE with pos_ids ([#609](#609)) ([ff05155](ff05155)) * improve plan performance by using non-blocking memcpy ([#547](#547)) ([41ebe6d](41ebe6d)) * reduce the read and write of shared memory in the FusedAddRMSNormKernel ([#592](#592)) ([2043ca2](2043ca2)) * reduce total_num_tiles_q by one ([#644](#644)) ([553ace5](553ace5)) * remove unnecessary contiguous operation in block sparse attention ([#561](#561)) ([7a7ad46](7a7ad46)) * speedup jit compilation of prefill attention kernels ([#632](#632)) ([a059586](a059586)) * use cuda-core implemention for io-bound block-sparse attention ([#560](#560)) ([3fbf028](3fbf028)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Zihao Ye <[email protected]>

Abatom added 2 commits November 6, 2024 14:30

Improve Precision

bcaea47

1e-2 -> 1e-3

26c0119

yzh119 reviewed Nov 6, 2024

View reviewed changes

include/flashinfer/norm.cuh Outdated Show resolved Hide resolved

Abatom added 2 commits November 6, 2024 19:01

use shared memory

6e384c3

use shared memory

bef7a5a

Abatom requested a review from yzh119 November 6, 2024 11:21

yzh119 approved these changes Nov 6, 2024

View reviewed changes

yzh119 merged commit c7dc921 into flashinfer-ai:main Nov 6, 2024

github-actions bot mentioned this pull request Nov 6, 2024

chore(main): release 0.2.0 #476

Merged

Abatom mentioned this pull request Nov 8, 2024

perf: reduce the read and write of shared memory in the FusedAddRMSNormKernel #592

Merged

yzh119 mentioned this pull request Nov 24, 2024

misc: remove duplicate norm cuda kernels #631

Merged

yzh119 mentioned this pull request Dec 1, 2024

[Bug] flashinfer's RMSNorm implementation causes precision differences in model outputs compared to the HuggingFace implementation sgl-project/sglang#2258

Closed

5 tasks

yzh119 mentioned this pull request Dec 25, 2024

Different sequence numbers calculate inconsistent results #696

Open

github-actions bot mentioned this pull request Dec 25, 2024

chore(main): release 0.3.0 #698

Closed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Improve the precision of the FusedAddRMSNormKernel function #587

Improve the precision of the FusedAddRMSNormKernel function #587

Abatom commented Nov 6, 2024 •

edited

Loading

Abatom commented Nov 6, 2024

yzh119 left a comment •

edited

Loading

zhyncs commented Nov 6, 2024

yzh119 commented Nov 6, 2024

Abatom commented Nov 6, 2024

yzh119 left a comment

Improve the precision of the FusedAddRMSNormKernel function #587

Improve the precision of the FusedAddRMSNormKernel function #587

Conversation

Abatom commented Nov 6, 2024 • edited Loading

Abatom commented Nov 6, 2024

yzh119 left a comment • edited Loading

Choose a reason for hiding this comment

zhyncs commented Nov 6, 2024

yzh119 commented Nov 6, 2024

Abatom commented Nov 6, 2024

yzh119 left a comment

Choose a reason for hiding this comment

Abatom commented Nov 6, 2024 •

edited

Loading

yzh119 left a comment •

edited

Loading