Hidden-state recurrence · Depth allocation · Length generalization

Recurrent Looped
Transformer

Recurrent computation across prompt and response.
Six algorithmic tasks, encoder–decoder depth allocation, and feedback-interval controls.

Yifan Zhang1  ·  Jichen Feng2  ·  Shihan Qin2
1Princeton University  ·  2University of Pennsylvania
arXiv:2610.07591  ·  Submitted October 6, 2026  ·  Updated October 7, 2026
Hidden-state feedbackFeedback intervalThree-seed experiments

Abstract

State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer over three seeds. Trained on at most 40 bits, two RLT splits generalize parity to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based \(S_5\) permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer, and accuracy increases with decoder depth. On modular arithmetic beyond the training lengths, RLT reaches up to 93% versus 33% for the Transformer. Ablations show that these gains depend on the feedback: removing it drops parity and swap-based \(S_5\) to chance at every split. Updating the feedback once per four-token chunk lets known tokens in a chunk run in parallel and keeps 64-bit parity at 99%, while permutation tracking depends on per-token feedback: chunking lowers length-64 swap-based \(S_5\) from 100% to 20%.

Architecture and execution

01 / REASONING

Depth that grows with the sequence.

Each token extends the recurrent path through the full decoder. After \(t\) tokens, that path traverses \(tL_D\) decoder blocks while the per-token block count stays fixed.

02 / HARDWARE

Parallel work around a recurrent core.

Batch known-token encoder work and independent decoder updates. Reuse weights and memory, and checkpoint activations while preserving the reference computation.

03 / RL

One transition from sampling to replay.

Rebuild the full history under current parameters, including prompt states and decoder SWA KV. Keep recorded behavior probabilities tied to the actual sampler.

The complete state matters.

RLT state updates across prompt and response.
The merge reads the previous final output; each SWA layer reads its own cache. Global memory is restricted to the current prefix.
\[H_t=(s_t,C_t^D),\qquad H_0=(s_\star,\varnothing).\]
\[(s_t,C_t^D)=D_\phi\!\left(\operatorname{Merge}(e_t,s_{t-1});M_{\le t},C_{t-1}^D,t\right).\]
\[p_\Theta(x_{t+1}\mid x_{1:t})=\operatorname{softmax}\!\left(W_o\operatorname{RMSNorm}_o(s_t)\right)_{x_{t+1}}.\]

Here \(M_{\le t}\) is global encoder memory, \(s_t\) is the recurrent output, and \(C_t^D\) contains layerwise decoder KV. A SWA window of \(W\) includes the current token and retains at most \(W-1\) historical entries for the next update.

Compatible encoder and decoder attention and FFN weights can be shared. A tied 48+48 layout illustrates this option in the report; the experiments below use untied eight- and sixteen-layer layouts.

Detailed RLT architecture
Detailed RLT architecture with global KV, per-layer SWA, gated merge and final-output feedback.
With one memory group (G = 1), all decoder layers read the same projected global KV using their own queries. Vector PDF ↗

Feedback interval: RLT-1, RLT-2 and RLT-0

RLT-1 and its two controls differ in one quantity, the feedback interval \(B\): the number of tokens between updates of the state that enters the decoder input. RLT-1 updates this state at every token (\(B=1\)), RLT-2 once per chunk of \(B\) tokens, and RLT-0 never (\(B=\infty\)). All three keep the causal encoder, encoder memory, decoder SWA, prefix-restricted memory attention and readout of RLT-1.

RLT-2 holds the feedback state fixed within a chunk and replaces it with the chunk's last decoder output at a complete boundary. Each position merges its encoder output with this boundary state through its own gate, and known positions run together within each decoder layer under causal SWA and prefix-restricted encoder memory. Boundaries are anchored at BOS, and a partial chunk keeps the previous boundary state, so moving the prompt–response split does not change the computation. With \(B=1\), RLT-2 reduces exactly to RLT-1.

RLT-2 architecture for one chunk: each position gates the same boundary state into its decoder input.
Every output predicts the next token; only the last output of a complete chunk becomes the next feedback state. Vector PDF ↗

RLT-0 is the \(B=\infty\) limit. It removes the feedback path together with the learned initial state, state normalization, gate and feedback projection, and feeds \(z_t^0=e_t\) to the decoder. At each eight-layer split, it has 787,968 fewer parameters than RLT-1.

Architecture of RLT-0
RLT without hidden-state feedback: direct encoder input, shared global KV and layerwise SWA.
Each layer processes known positions together under causal masks. Layers follow depth order. Vector PDF ↗

For a known prefix of length \(T\), the decoder needs \(\lceil T/B\rceil L_D\) sequential block stages: \(TL_D\) for RLT-1 and \(L_D\) for RLT-0. All three evaluate \(TL_D\) decoder blocks and generate one token per step.

RLT-1, RLT-2 and RLT-0 decoder schedules for eight known tokens.
Each shaded group is one layerwise decoder batch. Vector PDF ↗

Because RLT-2 uses the same parameters for every \(B\), the chunk size can also vary during training: pretraining can start with large chunks for parallelism, and mid- and post-training can reduce \(B\) toward 1. The experiments train each chunk size from scratch.

Results

Depth allocation · Feedback frequency · CPU step times · Addition and standard S5 · Complete six-task study · Sixteen-layer parity

The experiments ask how to allocate eight layers between the encoder and decoder, and how often to feed back the final decoder state. The six algorithmic tasks are addition, parity, modular arithmetic (mod 5) with and without brackets, and standard and swaps-based S5 permutation tracking. Every eight-layer comparison uses initialization seeds 42, 43 and 44 with shared training data and test sets, and reports mean ± sample standard deviation (SD, n = 3).

Models have width 512, FFN width 1,365 and four attention heads. RLT-1 uses an SWA window of eight, one encoder-memory group, feedback scale 0.1 and TBPTT 128, which covers every training sequence. In these comparisons, parity, addition and S5 train for 2,000 steps and mod 5 for 5,000. RLT-1 has 26.10–28.73M parameters and Transformer 8 has 25.31M, so equal layer counts do not match parameters or compute. Length generalization uses each run's checkpoint with the lowest in-distribution validation loss, taking the earliest step on ties, and evaluates 1,024 shared examples per length (256 operand pairs for addition).

Depth allocation and length generalization

Length generalization on parity, swaps-S5 and both mod-5 tasks for all five RLT-1 splits and Transformer 8, with three-seed error bars.
Gray regions mark training lengths, dotted lines mark uniform-prediction accuracy, and error bars show untrimmed sample SD. Parity and swaps-S5 use 2,000-step runs; both mod-5 tasks use 5,000-step runs for every model. PDF ↗ · SVG ↗
  • Parity: after training on at most 40 bits, 5+3 and 7+1 reach 100 ± 0% at 256 bits in all three seeds; Transformer 8 reaches 50.07 ± 1.63%. At step 500, 6+2 already reaches 99.44 ± 0.98% validation accuracy, versus 48.48 ± 0.53% for the Transformer.
  • Swaps-S5 favors a larger decoder. At 256 operations, eight times the training length, 4+4 reaches 97.30 ± 2.76% final-state accuracy, versus 0.85 ± 0.30% for the Transformer; at 512 operations it still reaches 55.70 ± 25.78%. Splits 7+1 and 8+0 are near the uniform reference at 256 operations despite high training-length accuracy.
  • Mod 5: on flat expressions of length 63, 6+2 reaches 93.36 ± 5.69%, versus 33.20 ± 2.33% for the Transformer. On bracketed expressions of length 64, 5+3 reaches 67.97 ± 2.91%, versus 46.71 ± 1.21%. Accuracy falls on longer expressions, and several flat mod-5 splits vary widely across seeds.
Parity accuracy at steps 500 and 2,000, with individual seeds, means and sample SD.
Both panels use the same 768 validation examples, after 256,000 and 1,024,000 training examples per seed. PDF ↗

Feedback frequency

RLT-0 and RLT-2 with four-token chunks (chunk4) were trained at every split with the data, optimizer and seeds of the RLT-1 runs. RLT-2 has the same parameters as RLT-1.

Feedback frequency at a fixed 4+4 split: RLT-1, RLT-2 chunk4, RLT-0 and Transformer 8 on parity, swaps-S5 and bracketed mod-5.
PDF ↗ · SVG ↗

At 64 bits, RLT-1 reaches 100 ± 0% parity accuracy and chunk4 98.99 ± 1.66%, while RLT-0 stays near chance at 50.23 ± 3.80%. Swaps-S5 is more sensitive to the interval: at 64 operations, RLT-1 reaches 100 ± 0% and chunk4 19.60 ± 6.54%, and RLT-0 and the Transformer are near the 1/120 uniform reference. Relative to RLT-1, four-token chunks lose about one percentage point on 64-bit parity and about 80 points on 64-operation swaps-S5.

The effect of chunking also depends on the split. Chunk4 4+4 retains 69.34 ± 24.17% parity accuracy at 256 bits, whereas chunk4 7+1 and 8+0 are near chance at 64 bits. On swaps-S5, chunk4 5+3 reaches 79.92 ± 10.89% at 48 operations and 50.16 ± 18.77% at 64.

Best-checkpoint test accuracy (%) at 4+4, mean ± sample SD over three seeds
Task / lengthRLT-1RLT-0RLT-2 chunk4Transformer 8
Parity / 64100.00 ± 0.0050.23 ± 3.8098.99 ± 1.6648.47 ± 1.18
S5 swaps / 48100.00 ± 0.007.26 ± 3.1768.29 ± 8.7522.10 ± 2.76
S5 standard / 321.53 ± 0.461.20 ± 0.721.37 ± 0.870.91 ± 0.15
Mod 5, flat / 6344.43 ± 43.8121.84 ± 2.1750.78 ± 38.0233.20 ± 2.33
Mod 5, brackets / 6464.10 ± 3.4939.45 ± 0.6162.76 ± 4.4246.71 ± 1.21
Parity length generalization for RLT-1, RLT-2 chunk4 and RLT-0 at every split.
Swaps-S5 length generalization for RLT-1, RLT-2 chunk4 and RLT-0 at every split.

Standard S5 · Flat mod 5 · Bracketed mod 5 · Validation at 4+4 · All variant figures

CPU training-step time

Seed-42 mean seconds per step at 4+4, over steps 1,001–2,000 with four CPU threads and FP32 on a shared cluster. The averages exclude validation, checkpoint writes and logging.

Model (4+4)Flat mod 5: seconds/stepBracketed mod 5: seconds/step
RLT-162.5954.09
RLT-2 chunk427.5223.81
RLT-014.5412.98

Chunk4 training steps run 2.27× as fast as RLT-1 steps, and RLT-0 steps 4.17–4.30× as fast. These are CPU training steps; wall-clock speed on other hardware depends on the implementation.

Addition, standard S5 and scope

All models reach 100% teacher-forced accuracy on 1–8-digit addition validation, but accuracy drops beyond eight digits; at 32 digits the eight-layer model means range from 14.89% to 16.84%. Standard S5 remains near the 1/120 reference across feedback variants. A three-seed ablation of the feedback scale finds no consistent effect on addition: across 63 comparisons with α = 0.1 at the same split and width, lowering α to 0.03 or 0.01 raises 37 means and lowers 26, and only three differences exceed both sample SDs (figure). These experiments use supervised training; RL performance is not evaluated.

Complete six-task study at 2,000 steps

The September 17, 2026 snapshot compares RLT-1 4+4, 5+3, 6+2, 7+1, 8+0 and Transformer 8 on all six tasks. All 108 runs completed 2,000 optimizer steps with initialization seeds 42, 43 and 44 for every task. The 8+0 variant has no decoder blocks but still applies the gated recurrent merge.

Validation accuracy, length generalization and diagnostics

Validation accuracy after 2,000 steps

Values are percentages, reported as mean ± sample SD across three initializations. Each run has consumed 1,024,000 training examples. Addition measures teacher-forced answer-token accuracy, including answer formatting and EOS, excluding prompt and padding positions. Parity and mod 5 score the final label; S5 scores the final state.

Task RLT-1 4+4 RLT-1 5+3 RLT-1 6+2 RLT-1 7+1 RLT-1 8+0 Transformer 8
Addition 100.00±0.00 100.00±0.00 100.00±0.00 100.00±0.00 100.00±0.00 100.00±0.00
Parity 100.00±0.00 100.00±0.00 100.00±0.00 100.00±0.00 98.83±1.92 94.84±3.43
Mod 5, no brackets 45.36±46.43 94.18±7.29 69.62±33.01 70.01±43.15 60.33±35.70 64.02±37.64
Mod 5, brackets 70.53±10.75 74.35±6.71 74.87±3.72 74.78±3.01 75.17±6.78 73.87±9.07
S5, swaps 100.00±0.00 100.00±0.00 100.00±0.00 99.35±0.81 99.61±0.68 99.09±0.23
S5, standard 0.78±0.68 2.47±1.26 2.21±1.48 2.08±0.98 1.82±1.19 0.52±0.23
Validation accuracy on six tasks, mean and sample SD across three initialization seeds.
Bands show sample SD, clipped to the accuracy range. Standard S5 uses a narrower vertical scale. PDF ↗

At 2,000 steps, flat mod 5 varies strongly with initialization: the three 4+4 seeds score 17.58%, 19.53% and 98.96%. Bracketed mod-5 means are closer, spanning 70.53–75.17% for RLT-1 versus 73.87 ± 9.07% for the Transformer. The 5,000-step runs above train both mod-5 tasks longer.

Length generalization

The evaluation covers 846 model–seed–length combinations, summarized as 282 three-seed means and sample SDs. At every task and length, all models and seeds receive the same 256 addition pairs or 1,024 formal-task sequences.

Length generalization on all six tasks with three-seed error bars.
Gray regions mark training lengths. Addition uses teacher-forced answer tokens; other tasks use final labels or states. PDF ↗

Accuracy at the longest tested length, in percent. Addition lengths count digits per operand; formal-task lengths count input symbols or operations, excluding boundary markers.

Task (test length) RLT-1 4+4 RLT-1 5+3 RLT-1 6+2 RLT-1 7+1 RLT-1 8+0 Transformer 8
Addition (32) 15.30±0.44 14.89±2.27 15.74±1.58 15.59±1.57 15.31±0.93 16.84±1.45
Parity (256) 66.76±28.78 100.00±0.00 84.05±27.63 100.00±0.00 68.91±27.39 50.07±1.63
Mod 5, no brackets (255) 18.00±0.39 20.57±0.62 21.42±0.49 19.34±0.54 20.44±0.91 20.35±2.01
Mod 5, brackets (256) 25.20±1.71 25.81±4.34 21.58±1.86 22.04±1.72 22.30±1.13 25.07±2.05
S5, standard (512) 1.24±0.30 1.43±0.60 1.30±0.31 0.88±0.20 0.72±0.31 0.81±0.06
S5, swaps (512) 55.70±25.78 34.86±6.10 22.14±20.95 0.85±0.06 0.85±0.20 0.85±0.30
S5 length generalization under prefix-token, final-state and whole-sequence scoring.
Whole-sequence accuracy requires every prefix prediction to be correct. Each panel labels its vertical scale. PDF ↗

Training loss, individual parity seeds, fixed-length token accuracy, and token-level length generalization provide additional diagnostics. Depth-eight figures and protocol

Sixteen-layer parity

A separate seed-42 series trains RLT-1 8+8 through 16+0 and Transformer 16 on parity lengths 3–40 with global batch 1,024. All ten runs completed 2,000 steps. Comparing each run's best in-distribution checkpoint with its step-2,000 checkpoint on the same 1,024 examples per length leaves 74 of 80 accuracies unchanged. 8+8, 9+7, 11+5 and 16+0 retain 100% at 256 bits at both checkpoints, compared with 49.41% for Transformer 16.

Sixteen-layer parity: best versus step-2,000 checkpoints, seed 42.
Accuracy (%) at 256 bits, seed 42
ModelBest stepBest checkpointStep 2,000
RLT-1 8+81,400100.00100.00
RLT-1 9+71,300100.00100.00
RLT-1 10+61,80062.7066.11
RLT-1 11+51,300100.00100.00
RLT-1 12+41,00098.8398.83
RLT-1 13+32,00099.4199.41
RLT-1 14+21,50063.5762.11
RLT-1 15+11,10094.1494.82
RLT-1 16+01,900100.00100.00
Transformer 162,00049.4149.41

Independent community experiments

Earlier state-tracking results contributed by @AradhyeAgarwal use a separate implementation with about 79K parameters, three seeds and training length 32. Evaluation uses 2,048 test programs per task and length, extending to 128 operations.

Community length-extrapolation results
Independent community comparison on parity and five-state transitions across sequence lengths.
Points show seed means and whiskers show seed minima and maxima. RLT reaches 60.8% parity accuracy and 20.7% five-state accuracy at 128 operations; chance levels are 50% and 20%. Parameter and data budgets were matched; FLOPs were not.

Per-length tables and measurement notes

One execution across training and inference

Known tokens can be encoded in a causal batch. Decoder updates still proceed in token order, constructing both recurrent outputs and decoder SWA caches.

Reference execution schedules
ModeEncoderDecoder
Prompt prefillCausal batchUpdate complete state through every prompt token.
GenerationIncrementalSample from the preceding state, then consume each token exactly once.
PretrainingCausal batchFull BPTT over all valid next-token targets.
SFTCausal batchAssistant-target loss; all context tokens update differentiable state.
RL replayRebuild with current weightsReplay the complete history and SWA caches; score each action before consuming it.

Forward consistency and complete gradients are separate requirements. Full BPTT includes paths through recurrent outputs, decoder KV, and encoder memory. Detaching any of these changes the gradient. Parameter updates invalidate old caches for exact current-policy replay.

Behavior log-probabilities must describe the actual sampling distribution. Exact importance sampling additionally requires support coverage. Shared transitions remove structural prompt-boundary mismatch; numerical kernel parity and off-policy estimation remain separate concerns.

For the execution-level distinction, see the prefill–decode kernel mismatch note.

Read the paper on arXiv ↗

Citation

If you find this work useful, please cite:

@misc{zhang2026recurrentlooped,
  title         = {Recurrent Looped Transformer},
  author        = {Zhang, Yifan and Feng, Jichen and Qin, Shihan},
  year          = {2026},
  eprint        = {2610.07591},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2610.07591}
}