Skip to content

[Question] Support head_dim=256 for attention ops (Qwen3.5/3.6-style GDN-hybrid models) ? #70

Description

@yaofeino1
  • The attention ops currently support only head_dim = 80 / 128.Could head_dim=256 be added?

  • Recent GDN-hybrid models (Qwen3.5 / 3.6) use head_dim=256 in their full-attention layers, so this would unblock those configs.

  • One question: since GDN (linear-attention) layers dominate these models,how much end-to-end latency benefit can we realistically expect from optimizing the full-attention op here? Has anyone benchmarked hpc_ops attention on a Qwen3.5/3.6-class model?

  • Happy to help test if a branch is available.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions