Hello,
Thank you for implementing this feature. I did some quick tests but don't see much improvement from orthogonalizing heads independently (needs verification). One change is in the adjust_lr of q/k/v, which I think is scaled smaller by 1/√H when using num_heads=12. I don't quite understand the implications. Allowing heads to orthogonalize independently probably generates larger updates, but could reducing lr by 3.46x in my 12-head model be over-compensating for that? Is this something I have to tune?
Would appreciate your insights. Thanks!
Hello,
Thank you for implementing this feature. I did some quick tests but don't see much improvement from orthogonalizing heads independently (needs verification). One change is in the adjust_lr of q/k/v, which I think is scaled smaller by 1/√H when using num_heads=12. I don't quite understand the implications. Allowing heads to orthogonalize independently probably generates larger updates, but could reducing lr by 3.46x in my 12-head model be over-compensating for that? Is this something I have to tune?
Would appreciate your insights. Thanks!