You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The attention ops currently support only head_dim = 80 / 128.Could head_dim=256 be added?
Recent GDN-hybrid models (Qwen3.5 / 3.6) use head_dim=256 in their full-attention layers, so this would unblock those configs.
One question: since GDN (linear-attention) layers dominate these models,how much end-to-end latency benefit can we realistically expect from optimizing the full-attention op here? Has anyone benchmarked hpc_ops attention on a Qwen3.5/3.6-class model?
The attention ops currently support only head_dim = 80 / 128.Could head_dim=256 be added?
Recent GDN-hybrid models (Qwen3.5 / 3.6) use head_dim=256 in their full-attention layers, so this would unblock those configs.
One question: since GDN (linear-attention) layers dominate these models,how much end-to-end latency benefit can we realistically expect from optimizing the full-attention op here? Has anyone benchmarked hpc_ops attention on a Qwen3.5/3.6-class model?
Happy to help test if a branch is available.