Implement custom kernel for LLaMA rotary embedding by WoosukKwon · Pull Request #14 · vllm-project/vllm

WoosukKwon · 2023-03-30T10:27:36Z

This PR implements a custom CUDA kernel for rotary embedding, which is used in LLaMA. The kernel is responsible for the entire process of applying rotary embedding to query and key, and is thus much more efficient than the PyTorch implementation.

Tested models:

LLaMA-7B
LLaMA-13B

Tested GPUs:

A100

zhuohan123

LGTM!

csrc/pos_encoding_kernels.cu

Install NNCF

* remove JambaConfig and use official one from transformers * changes in Jamba modeling file to align with official HF format

enable fused topK_softmax kernel for hip path

0612 kernel of FP8 on A100

Summary: Add benchmarking scripts and utils. Things to note : - All files are stored in `neuralmagic` folder. - neuralmagic/benchmarks/scripts/* : Actual benchmarking scripts that interact with vllm engine. - neuralmagic/benchmarks/configs/* : JSON config files that define what benchmark commands to run. - neuralmagic/benchmarks/run_*.py : Scripts that consume some config file and run the benchmark scripts. - neuralmagic/tools : Add tools Testing: Local testing --------- Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com> Co-authored-by: rsnm2 <rshaw@neuralmagic.com>

a fix follow up [MRotaryEmbedding change](vllm-project@bf3b79e#diff-6bc44986c91bf0876240dec03d56c748403691c7fcd90f7a22e7affff7b033ecR839) Signed-off-by: z00897138 <zhaorifa@huawei.com> Co-authored-by: z00897138 <zhaorifa@huawei.com>

setup sparse attention backend

…pSeek-v2 (vllm-project#28101) (vllm-project#14) Signed-off-by: Kunshang Ji <kunshang.ji@intel.com> Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn> Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>

… update perf measurement to decode multiple tokens Signed-off-by: Salar Hosseini <skhorasgani@tenstorrent.com>

[Model] Add end2end example and documentation for qwen2.5-omni

…-project#7) Adds five new fields to LoRAConfig in vllm/config/lora.py to support runtime dynamic resizing of GPU LoRA adapter slots: - min_loras (int, ge=1): floor for dynamic slot shrinking - dynamic_lora_slots (bool): enables automatic watermark-driven scaling - lora_mem_high_watermark (float, 0<x<1): scale-down threshold - lora_mem_low_watermark (float, 0<x<1): scale-up threshold - lora_slot_resize_cooldown_s (float, ge=0): anti-thrash cooldown Cross-field validation added to _validate_lora_config(): - min_loras <= max_loras (when dynamic_lora_slots=True) - lora_mem_low_watermark < lora_mem_high_watermark (when dynamic=True) Field-level bounds (ge/gt/lt) enforced by Pydantic at construction time. dynamic_lora_slots added to compute_hash() as it affects the CudaGraph specialization path (disables LoRA cudagraph when True, see issue vllm-project#14). All new fields default to safe values so existing configs are unaffected when dynamic_lora_slots=False (the default). Includes 16 unit tests in tests/lora/test_lora_config_dynamic.py covering defaults, valid configs, all validation error paths, and compute_hash() behavior. Closes vllm-project#7 Closes vllm-project#18 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Signed-off-by: Chen Wang <Chen.Wang1@ibm.com>

…e_hash - Clarify min_loras docstring: <= max_loras is only enforced when dynamic_lora_slots=True, not unconditionally. - Clarify dynamic_lora_slots docstring: remove reference to unimplemented POST /v1/scale_max_loras endpoint; note operator-triggered scaling will be handled via plugin (issue vllm-project#16). - Fix compute_hash() comment with TODO(vllm-project#14) reference. - Add specialize_active_lora to compute_hash() factors — it controls which CUDA graphs are captured and must be part of the computation graph hash. - Add test_compute_hash_differs_with_specialize_active_lora to cover above. Co-authored-by: Claude Signed-off-by: Chen Wang <Chen.Wang1@ibm.com>

…-section-12 Expand README Section 12 experimental results and add Section 13 Algorithmic Efficiency analysis

WoosukKwon added 7 commits March 30, 2023 07:01

Minor

512c7bf

Add test code for rotary embedding

56674f4

Minor

3533de0

Minor

8e0e6a4

Add rotary embedding kernel

3b6652a

Add test code for rotary embedding kernel

b29eb16

Implement Llama attention layer

7392665

WoosukKwon requested a review from zhuohan123 March 30, 2023 10:29

Minor fix in comment

eef11ba

WoosukKwon changed the title ~~Add custom kernel for rotary embedding~~ Implement custom kernel for LLaMA rotary embedding Mar 30, 2023

zhuohan123 approved these changes Mar 30, 2023

View reviewed changes

csrc/pos_encoding_kernels.cu Show resolved Hide resolved

Test more head sizes

1a26188

WoosukKwon merged commit 88c0268 into main Mar 30, 2023

WoosukKwon deleted the rotary-embedding branch March 30, 2023 18:04

bigPYJ1151 added a commit to bigPYJ1151/vllm that referenced this pull request Sep 12, 2023

Add multi-attention op. (vllm-project#14)

824dfc9

shanshanpt mentioned this pull request Nov 17, 2023

Run long conetxt error : CUDA error: an illegal memory access was encountered #1700

Closed

junior-zsy mentioned this pull request Nov 20, 2023

Error with 32k Long Text in chatglm2-6b-32k Model #1725

Closed

hongxiayang pushed a commit to hongxiayang/vllm that referenced this pull request Feb 13, 2024

Implement custom kernel for LLaMA rotary embedding (vllm-project#14)

cb12020

luo-cheng2021 pushed a commit to luo-cheng2021/vllm that referenced this pull request Mar 25, 2024

Merge pull request vllm-project#14 from ilya-lavrenov/install-nncf

05b9161

Install NNCF

mzusman pushed a commit to mzusman/vllm that referenced this pull request May 6, 2024

Jamba official hf (vllm-project#14)

988718e

* remove JambaConfig and use official one from transformers * changes in Jamba modeling file to align with official HF format

fxmarty pushed a commit to fxmarty/vllm-public that referenced this pull request May 31, 2024

Merge pull request vllm-project#14 from ROCm/fused_topK_softmax

e3ae076

enable fused topK_softmax kernel for hip path

yuhuixu1993 mentioned this pull request Jun 2, 2024

[Bug]: loading squeezellm model #5190

Closed

ykim362 pushed a commit to ykim362/vllm that referenced this pull request Jun 17, 2024

Merge pull request vllm-project#14 from wenxcs/wenxh/fp8-on-a100-v5-pr

b28848e

0612 kernel of FP8 on A100

alixiaodi mentioned this pull request Aug 2, 2024

[Bug]: #7072

Closed

SpaceHunterInf mentioned this pull request Sep 30, 2024

[Bug]: Bus error (core dumped) #8974

Closed

1 task

hao-cold mentioned this pull request May 13, 2025

[Bug]: CUDA error: an illegal instruction was encountered #18045

Closed

1 task

markmc mentioned this pull request May 21, 2025

[Bug][Failing Test]: Distributed Comm Ops - distributed/test_shm_broadcast.py #18492

Closed

1 task

zerosurplus mentioned this pull request Jun 16, 2025

[Bug]: torch.distributed.DistNetworkError: The client socket has timed out after 600000ms while trying to connect to (172.17.0.9, 46229). #19670

Open

1 task

xiaomofang mentioned this pull request Jul 31, 2025

[Bug]: There is an issue with speculative inference in Eagle mode, where the context length of vLLM inference is constrained by the draft model. #21986

Closed

1 task

heheda12345 pushed a commit to heheda12345/vllm that referenced this pull request Sep 29, 2025

Merge pull request vllm-project#14 from vllm-model-0920/mla_backend

446c0de

setup sparse attention backend

Michel-debug mentioned this pull request Oct 23, 2025

[Bug]: qwen3-vl-2b after ms-swift fine-tuning lance errors #27405

Closed

1 task

acodercat mentioned this pull request Nov 10, 2025

[Bugfix] Add strong reference to CUDA pluggable allocator callbacks #23477

Merged

4 tasks

sravan500 mentioned this pull request Nov 25, 2025

[Bug]: vllm/vllm-openai:v0.11.0 deployment --quantization fp8 throws cuda and tensor errors #29374

Closed

1 task

iwooook pushed a commit to moreh-dev/vllm that referenced this pull request Nov 29, 2025

vllm-project#14: Add trace_mode option to TTWorker and TTModelRunner,…

18f4c6e

… update perf measurement to decode multiple tokens Signed-off-by: Salar Hosseini <skhorasgani@tenstorrent.com>

tjtanaa pushed a commit to tjtanaa/vllm that referenced this pull request Jan 29, 2026

Merge pull request vllm-project#14 from Gaohan123/end2end_example

d41d3e4

[Model] Add end2end example and documentation for qwen2.5-omni

Lrcx mentioned this pull request Jan 29, 2026

[Bug]: Crash when using presence_penalty with Qwen3-VL in v0.11.0 #33338

Open

1 task

HervorTao mentioned this pull request Feb 3, 2026

[Bug]: [CPU Backend] AttributeError: '_OpNamespace' '_C_utils' object has no attribute 'init_cpu_threads_env' #33675

Closed

1 task

LironKesem mentioned this pull request Mar 12, 2026

[Bug] DGX Spark (sm_121): CUTLASS can_implement() rejects sm_120f binaries #36835

Closed

1 task

mahaocong90 mentioned this pull request Mar 17, 2026

[Bug]: QWEN 3.5-397B-A17B report "RPC call to sample_tokens timed out" #37250

Closed

1 task

Copilot AI mentioned this pull request Mar 20, 2026

Fix XPU segfault when tensor_parallel_size exceeds available devices hongbolv/vllm#5

Closed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Implement custom kernel for LLaMA rotary embedding#14

Implement custom kernel for LLaMA rotary embedding#14
WoosukKwon merged 9 commits intomainfrom
rotary-embedding

WoosukKwon commented Mar 30, 2023 •

edited

Loading

Uh oh!

zhuohan123 left a comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Uh oh!

Conversation

WoosukKwon commented Mar 30, 2023 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

zhuohan123 left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

WoosukKwon commented Mar 30, 2023 •

edited

Loading