Skip to content

fix(turbomind): restore INT8 KV quant-param offset in block layout - #4764

Merged
lvhan028 merged 1 commit into
InternLM:mainfrom
lzhangzz:fix-int8-kv
Jul 20, 2026
Merged

fix(turbomind): restore INT8 KV quant-param offset in block layout#4764
lvhan028 merged 1 commit into
InternLM:mainfrom
lzhangzz:fix-int8-kv

Conversation

@lzhangzz

Copy link
Copy Markdown
Collaborator

No description provided.

After InternLM#4717 flattened per-layer layout to a byte offset, k_param() dropped
the data-region base (old layer_param = layer_data + head_data(head_num)),
so INT8 scale/zero writes overlapped K/V data and crashed (often surfacing
later in MoE). quant_policy=0 was unaffected because params are unused.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes the Turbomind attention KV-cache block memory layout so that INT8 (quantized) KV quantization parameters (scale/zero) are placed after the complete K/V data region for all heads, restoring the intended offset behavior.

Changes:

  • Adjust Layout::k_param() to add the base offset for the full K/V data region (head_data(head_num)) before indexing into per-head/per-token parameter storage.
  • Add an inline comment clarifying the intended layout and referencing the pre-#4717 behavior.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

@lvhan028
lvhan028 merged commit bfb549b into InternLM:main Jul 20, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants