A CONCEPT · PROJECTED FROM THE VINDEX KNOWLEDGE GRAPH
KDA — KIMI DELTA ATTENTION
fifteen operands, per-channel forgetting
A delta-attention operator with per-channel gating — a linear-attention family carried as its own execution surface, with its own declared geometry.
Fifteen operand roles: the Q, K and V projections and their short convolutions, the F-A/F-B and G-A/G-B gate pairs, a B projection, A_log and dt_bias, the output norm and the out-projection — plus kda_gate_lower_bound on the surface. Real stacks interleave it with full attention, and the interleave is read from the per-layer policy table, never from layer arithmetic.
IN THE GRAPH
kda — kimi delta attention— belongs to →linear attention
kda — kimi delta attention— sibling of →mla — multi-head latent attention
kda — kimi delta attention— carries →recurrent state
kda — kimi delta attention— governed by →hybrid attention policy