Deep Delta Learning

A residual update that reads the state along a learned direction, compares the readout with a learned target, and writes back a gated rank-1 correction.

Yifan Zhang1  ·  Yifeng Liu2  ·  Mengdi Wang1  ·  Quanquan Gu2

1Princeton University  •  2UCLA  •  January 1, 2026

Residual Connections Delta Rule Transformers Language Modeling

Abstract

Transformer residual streams evolve through additive updates. A sufficiently expressive residual block can represent content replacement, but standard architectures do not parameterize reading, comparison, and replacement as an explicit residual operation. We introduce Deep Delta Learning (DDL), a structured residual update that keeps the identity path and adds target-seeking edits to the residual state. Each layer reads the current state along a learned direction, compares the readout with a learned target, and writes back a gated rank-1 correction along the same direction. Closing the gate recovers the identity map; a unit gate exactly overwrites the selected residual readout.

We instantiate DDL with both scalar and expanded residual states. The expanded state stores multiple persistent value channels while attention and MLP computation stay at the original model width, so residual-state capacity can grow without widening the backbone. In controlled single-run pretraining comparisons at two scales, DDL improves validation loss and average one-shot downstream accuracy over additive residual baselines, with lower throughput in every measured configuration and higher peak memory for expanded states. These results suggest that depth-wise delta-rule updates provide a useful inductive bias for managing Transformer residual streams.

Project Repository Read the Paper

The Delta Residual Update

For a residual state $\Xb_l \in \RR^{d \times d_v}$, DDL computes

$$ \Xb_{l+1} = \Xb_l + \beta_l \kb_l \bigl(\vb_l^\top - \kb_l^\top \Xb_l\bigr) = (\Ib - \beta_l \kb_l \kb_l^\top)\Xb_l + \beta_l \kb_l \vb_l^\top, $$
  • $\kb_l \in \RR^d$ is a unit read/write direction: the normalized output of the attention or MLP sublayer;
  • $\vb_l \in \RR^{d_v}$ is the target readout, produced by a lightweight branch;
  • $\beta_l = 2\sigma(\cdot) \in (0, 2)$ is a gate shared by erasure and writing.
Deep Delta Learning overview. (a) The residual rewrite: read, compare, and write steps next to the identity path. (b) A DDL Transformer sublayer: compress, normalize, attention or MLP, and rewrite, with value and gate branches.
Figure 1: Deep Delta Learning overview. (a) Read the selected residual content, compare it with a target, and add a gated rank-1 correction to the identity path. (b) In a DDL Transformer sublayer, the attention or MLP output gives the direction; lightweight branches produce the target from the sublayer input $\mathbf{x}_l^{\mathrm{in}}$ and the gate from the normalized context $\mathbf{c}_l$. The expanded residual state persists across sublayers.

With $d_v = 1$ the state is the ordinary residual vector. With $d_v > 1$ it stores several value channels, and a learned compressor gives each attention or MLP block a width-$d$ input, so the expensive sublayers are not widened.

The update is still additive, so DDL does not enlarge the function class; it makes the edit target-seeking. After the update, the readout error along $\kb_l$ is multiplied by $1 - \beta_l$: closing the gate leaves the state unchanged, and $\beta_l = 1$ removes the error exactly.

Spectral Analysis

For a given direction $\kb$ and gate $\beta$, the shortcut $\Ab = \Ib - \beta \kb\kb^\top$ has eigenvalue $1$ on $\kb^\perp$ (multiplicity $d-1$) and $1-\beta$ along $\kb$. This gives three local regimes:

Regime Gate Eigenvalue along $\kb$ Effect on the readout $\kb^\top \Xb$
Skip $\beta \approx 0$ $\approx 1$ The update approaches the identity.
Target match $\beta = 1$ $0$ The readout is replaced exactly by $\vb^\top$.
Over-relaxed $1 < \beta < 2$ $1-\beta < 0$ The readout crosses the target. As $\beta \to 2$, the shortcut (not the full update) approaches the Householder reflector $\Ib - 2\kb\kb^\top$.

This analysis describes the operator for a given direction and gate. It does not show that learned directions correspond to human-readable features.

Depth-Wise Delta Rule

DeltaNet applies the delta rule over time to update a memory matrix. DDL applies the same erase/write update over network depth, as the residual interface between Transformer sublayers:

$$ \Xb_{l+1} = \Xb_l + \beta_l \kb_l \bigl(\underbrace{\vb_l^\top}_{\text{target}} - \underbrace{\kb_l^\top \Xb_l}_{\text{readout}}\bigr) $$

The rule itself is prior work; DDL's contribution is its depth-wise use and analysis.

Results

Decoder-only models (~124M and ~353M parameters) trained on FineWeb-Edu for 49.15B tokens. Validation loss and average one-shot accuracy (%) over eight benchmarks:

Model Small loss Small 1-shot Medium loss Medium 1-shot
Baseline2.854348.562.605353.96
DDL ($d_v=1$)2.848248.732.603954.69
DDL-TC w/o EC2.835548.912.592754.83
DDL-CC w/o EC2.832149.132.579054.92
DDL-TC2.829949.472.590554.86
DDL-CC2.832949.292.575855.14

TC compresses the expanded state over tokens, CC mixes its $d_v = 4$ value channels at each token, and EC initializes it with a convolution over token embeddings. Expanded-state variants cost throughput and memory: at the small scale, DDL-CC trains at 1158.0K tokens/s with 3.08 GB peak memory, against 1509.6K tokens/s and 2.94 GB for the baseline.

Each configuration was trained once with the same token budget, so these are point estimates, not compute-matched comparisons, and the expanded-state gains are not separated from the added residual capacity. The paper discusses these limits in detail.

Code

The repository contains PyTorch implementations of every reported variant. The reported runs use each file's default configuration. Each -accelerated file is a Triton implementation of the same model for CUDA GPUs; the reference files also run on CPU.

Paper nameModel file
DDL ($d_v=1$)model/DDL-vdim1-gpt-mha-rope-TC.py
DDL-TC w/o ECmodel/DDL-gpt-mha-rope-TC.py
DDL-TCmodel/DDL-gpt-mha-rope-TC-EC.py
DDL-CC w/o ECmodel/DDL-gpt-mha-rope-CC.py
DDL-CCmodel/DDL-gpt-mha-rope-CC-EC.py
import importlib
import torch

ddl = importlib.import_module("model.DDL-gpt-mha-rope-CC-EC")  # file names contain hyphens
config = ddl.GPTConfig()  # defaults: the ~124M model from the paper, d_v = 4
model = ddl.GPT(config)
input_ids = torch.randint(0, config.vocab_size, (1, 16))
out = model(input_ids=input_ids, labels=input_ids)
print(out.loss)

Citation

If you find this work useful, please cite:

@article{zhang2026deep,
   title   = {Deep Delta Learning},
   author  = {Zhang, Yifan and Liu, Yifeng and Wang, Mengdi and Gu, Quanquan},
   journal = {arXiv preprint arXiv:2601.00417},
   year    = {2026}
}