With raise of GDN and different sorts of attention meachanisms, those are much closer to lstm/recurrent architectures being very stateful rather than normal attention, we aim to explain explore common patters in lstm and GDN.
With raise of GDN and different sorts of attention meachanisms, those are much closer to lstm/recurrent architectures being very stateful rather than normal attention, we aim to explain explore common patters in lstm and GDN.
Project Details
Updated 07/04/26 · Provided via application · VerifiedIf you look closely on GDN from Qwen(Gated DaltaNet) attention and how it's actually implemented, you will notice it looks extremently like classic rnn's which motivates looking first at purely rnn's specific circuits and how they appear in gnd. The experiments might look like pertaining a very small couple layers lstm and gdn+mlp transformer blocks and reverse engineering certain circuits. I have a couple of concrete results already. the output is a paper + code, before the neurlips submission late September.
Theory of Impact
Updated 07/04/26 · By grantmaking.aiExisting interpretability tools do not natively expose rnn like circuitry and rely heavily on linearity of representations. This is not quite true in rnn's and the units invariant under model symmetries there are just coordinates. which makes circuitry more combinatorial like and less linear algebra like.
People
Updated 07/04/26 · By grantmaking.aiTeam Member
Private comment. Only shown to approved funders and grant reviewers.