For a residual state $\Xb_l \in \RR^{d \times d_v}$, DDL computes
- $\kb_l \in \RR^d$ is a unit read/write direction: the normalized output of the attention or MLP sublayer;
- $\vb_l \in \RR^{d_v}$ is the target readout, produced by a lightweight branch;
- $\beta_l = 2\sigma(\cdot) \in (0, 2)$ is a gate shared by erasure and writing.
With $d_v = 1$ the state is the ordinary residual vector. With $d_v > 1$ it stores several value channels, and a learned compressor gives each attention or MLP block a width-$d$ input, so the expensive sublayers are not widened.
The update is still additive, so DDL does not enlarge the function class; it makes the edit target-seeking. After the update, the readout error along $\kb_l$ is multiplied by $1 - \beta_l$: closing the gate leaves the state unchanged, and $\beta_l = 1$ removes the error exactly.