Sixteen tasks from the literature, one harness, three models, thirty-five dollars.
Mark Henry
For research I wanted a task where the model has to store intermediate state in its chain of thought (CoT) as it works. It's simple to test whether it's storing state in its CoT: if you blind the model to some of those intermediate tokens, the accuracy on the task will drop.
A good task has a high density of thinking tokens and fails when they go missing. This is the kind of task you want for experiments on faithfulness, steganography, or CoT monitoring.
There are a lot of candidate tasks in the literature and no comparison of them on this axis, so I made one. I collected twenty-five task designs from the scratchpad, state-tracking, and CoT-faithfulness literature, screened them on paper, implemented the sixteen task designs that survived the initial screening, and ran them through a battery of tests. The code is at github.com/mark-henry/stateful-tasks-review.
The models: Qwen3.5-9B with thinking disabled, Llama-3.3-70B, and DeepSeek-V4-Flash with thinking disabled, all served by together.ai at temperature 0. DeepSeek-V4-Pro was used once as a spot check. The roster and the survival rules were fixed before the runs, and are given in the methods section at the end.
If accuracy is the same regardless of whether thinking is on, the model isn't keeping anything in the CoT tokens. A task can be too easy at short lengths and impossible at long ones, so I ran each at a range of depths (difficulty levels, the number of steps a problem takes). A task passes on a model if, at some depth, thinking out loud lifts accuracy by at least 20 points and ends at 70% or better.
Twelve of the sixteen pass on at least one model, and seven pass on all three. In the table, a wide blue band shows improvement thanks to CoT. When both dots sit together near the left, the task is too hard either way; when they sit together near the right, too easy. A red band means thinking out loud made things worse, somehow.
accuracy without CoT with CoT the uplift 0.7, the minimum to pass (drawn solid where CoT falls short of it)
| task | Qwen3.5-9B | Llama-3.3-70B | DeepSeek-V4-Flash |
|---|---|---|---|
| Pass on all three models (7) | |||
| blocksworld | 0.64 → 0.92 at d2passes d2–d6 | 0.58 → 0.96 at d24passes d16–d24 | 0.64 → 0.90 at d24passes d24 |
| boolean expressions | 0.34 → 0.70 at d12passes d3–d12 | 0.62 → 0.88 at d4passes d4 | 0.56 → 0.96 at d8passes d4–d12 |
| cup shuffling | 0.22 → 0.98 at d12passes d2–d16 | 0.26 → 1.00 at d3passes d2–d12 | 0.22 → 1.00 at d3passes d2–d24 |
| Dyck | 0.20 → 0.88 at d8passes d4–d16 | 0.52 → 0.82 at d8passes d4–d8 | 0.18 → 0.86 at d24passes d4–d48 |
| nested arithmetic | 0.00 → 0.96 at d7passes d2–d20 | 0.02 → 0.96 at d7passes d2–d10 | 0.00 → 0.92 at d7passes d3–d14 |
| program trace† | 0.08 → 0.94 at d32passes d12–d48 | 0.04 → 1.00 at d48passes d12–d72 | 0.00 → 1.00 at d72passes d20–d72 |
| tag system | 0.04 → 0.90 at d2passes d2–d5 | 0.02 → 0.94 at d2passes d2–d4 | 0.00 → 1.00 at d8passes d2–d30 |
| Pass on two (2) | |||
| CRUXEval | 0.44 → 0.86 at d6passes d1–d12 | 0.26 → 0.56 at d15fails: acc below 0.7 | 0.50 → 0.86 at d6passes d1–d12 |
| entity tracking* (ours) | 0.40 → 0.82 at d8passes d4–d8 | 0.74 → 0.84 at d26fails: no gap | 0.60 → 0.98 at d36passes d8–d36 |
| Pass on DeepSeek-V4-Flash only; ergonomic versions below (3) | |||
| addition† | 0.00 → 0.00 at d24fails: no gap | 0.12 → 0.62 at d2fails: acc below 0.7 | 0.08 → 1.00 at d2passes d2 |
| random lookup table† | 0.12 → 0.44 at d2fails: acc below 0.7 | 0.12 → 0.52 at d4fails: acc below 0.7 | 0.14 → 1.00 at d24passes d2–d48 |
| S5 composition† | 0.02 → 0.06 at d2fails: no gap | 0.00 → 0.48 at d4fails: acc below 0.7 | 0.00 → 0.88 at d16passes d12–d16 |
| Fail on every model (4); all but multiplication get an ergonomic version below | |||
| cellular automaton* (ours) | 0.00 → 0.16 at d2fails: no gap | 0.00 → 0.04 at d2fails: no gap | 0.00 → 0.54 at d2fails: acc below 0.7 |
| Tower of Hanoi | 0.00 → 0.08 at d5fails: no gap | 0.04 → 0.14 at d2fails: no gap | 0.08 → 0.50 at d2fails: acc below 0.7 |
| multiplication | 0.04 → 0.08 at d30fails: no gap | 0.06 → 0.12 at d20fails: no gap | not run |
| 3SUM† | 0.42 → 0.22 at d2fails: no gap | 0.54 → 0.58 at d22fails: no gap | 0.44 → 0.28 at d2fails: no gap |
Two conventions used throughout.
Multiplication is dead, and addition nearly so, for the same reason: multi-digit addition and multiplication are solved in one forward pass by a 2026 9B model up to the widths where the published scratchpads were designed to help. Qwen3.5-9B multiplies 3×3-digit numbers without a scratchpad at 100%, and the Dziri et al. scratchpad only hurts. The same is true of the Nye et al. addition scratchpad on the 9B: 0.92 without it at two digits, 0.14 with it. There is nothing to blind.
Tower of Hanoi as published (the Apple formulation: produce the full move list) is dead on all three models, for a different reason: the published answer is the trace, so there is no "with CoT" condition to speak of. At two and five disks every model fails cleanly; from ten disks up the move lists overrun the token cap. 3SUM as published mostly fails to finish at all: between a third and four fifths of its traces hit the cap without an answer, even at 8,192 tokens on DeepSeek, and where a trace does finish it is at chance. Both get an ergonomic version below.
The 70B is not the strongest of the three. Qwen3.5-9B goes deeper than Llama-3.3-70B on 7 of 9 tasks, and DeepSeek goes at least as deep as the 70B on every task but blocksworld.
Four of the sixteen tasks come from papers that trained the model on the trace format (we've been denoting these tasks with a dagger †). So for each we wrote an "ergonomic" version which adapts the task to prompted models, writing the state out in full at every step and putting the prompt in plain language. With this, three of the four come back to life on both models. The fourth task, addition, is disqualified because the models no longer need CoT to accomplish it. Tower of Hanoi and the cellular automaton, which failed for other reasons, got ergonomic versions too; from here on, a row marked (ergonomic) uses that version.
| task | Qwen3.5-9B | DeepSeek-V4-Flash |
|---|---|---|
| Tower of Hanoiexecute k moves, report the pegs | as published0.00 → 0.08 at d5 · failsergonomic0.10 → 0.96 at d8 · passes | as published0.08 → 0.50 at d2 · failsergonomic0.02 → 1.00 at d32 · passes |
| cellular automaton*rev 2: one line per cell | our first format0.00 → 0.16 at d2 · failsergonomic0.02 → 0.59 at d1 · fails | our first format0.00 → 0.54 at d2 · failsergonomic0.02 → 0.78 at d1 · passes |
| cellular automaton*rev 3: named neighbours, running row | our first format0.00 → 0.16 at d2 · failsergonomic0.02 → 1.00 at d1 · passes | our first format0.00 → 0.54 at d2 · failsergonomic0.02 → 0.98 at d1 · passes |
| S5 composition† | as published0.02 → 0.06 at d2 · failsergonomic0.00 → 0.92 at d2 · passes | as published0.00 → 0.88 at d16 · passesergonomic0.00 → 1.00 at d4 · passes |
| random lookup table† | as published0.12 → 0.44 at d2 · failsergonomic0.04 → 1.00 at d8 · passes | as published0.14 → 1.00 at d24 · passesergonomic0.08 → 1.00 at d8 · passes |
| 3SUM† | as published0.42 → 0.22 at d2 · failsergonomic0.52 → 0.98 at d4 · passes | as published0.44 → 0.28 at d2 · failsergonomic0.48 → 1.00 at d4 · passes |
| addition† | as published0.00 → 0.00 at d24 · failsergonomic1.00 → 1.00 at d4 · fails | as published0.08 → 1.00 at d2 · passesergonomic1.00 → 1.00 at d6 · fails |
| task, depth | full trace → last step removed |
|---|---|
| Dyckd8 | 0.88 → 0.02 without the last step · 0.20 without thinking |
| tag systemd4 | 0.76 → 0.00 without the last step · 0.02 without thinking |
| 3SUM† (ergonomic)d8 | 0.92 → 0.46 without the last step · 0.48 without thinking |
| cellular automaton* (ergonomic)d3 | 0.82 → 0.00 without the last step · 0.00 without thinking |
| Tower of Hanoi (ergonomic)d16 | 0.86 → 0.00 without the last step · 0.00 without thinking |
| S5 composition† (ergonomic)d8 | 0.78 → 0.04 without the last step · 0.00 without thinking |
| random lookup table† (ergonomic)d24 | 1.00 → 0.10 without the last step · 0.06 without thinking |
| cup shufflingd12 | 0.98 → 0.40 without the last step · 0.26 without thinking |
| nested arithmeticd7 | 0.96 → 0.42 without the last step · 0.00 without thinking |
| program trace†d32 | 0.94 → 0.54 without the last two steps · 0.02 without thinking |
Nine of ten tasks fall to minimum accuracy if we remove even one step. Therefore the 9B cannot do one permutation, one swap, one rewrite, one lookup in its head for these, and the trace is load-bearing. The one exception is nested arithmetic, where the last step is a single combination of two named sub-results, which the model can manage 42% of the time.
| task, depth | propagation | what happens |
|---|---|---|
| random lookup table† (ergonomic)d24 | 0.99 | a wrong symbol is carried through to the answer |
| nested arithmeticd7 | 0.96 | a wrong sub-result is carried through to the answer |
| S5 composition† (ergonomic)d8 | 0.89 | a wrong permutation is carried through to the answer |
| cellular automaton* (ergonomic)d3 | 0.88 | a wrong row is carried through to the answer |
| cup shufflingd12 | 0.73 | after corruption, accuracy is at chance for three possible answers, so the mistake is carried through almost fully |
| program trace†d32 | 0.65 | corrupted values are often overwritten later |
| Dyckd8 | 0.53 | a wrong stack is often absorbed by later pops |
| tag systemd4 | 0.51 | partial; I did not look into why |
| Tower of Hanoi (ergonomic)d16 | 0.33 | model detects the illegal move and recomputes |
| 3SUM† (ergonomic)d8 | -0.11 | undefined: the uncorrupted continuation itself fails (below) |
This is the complement of knockout: instead of removing the state, corrupt it and see whether the model reads what it wrote. For four tasks the answer is yes, nearly always: the lookup table (0.99), nested arithmetic (0.96), permutation composition (0.89) and the cellular automaton (0.88). A perturbation in the state becomes an incorrect answer. Those four are the cleanest channels this review found: every step load-bearing, state written and read, no escape hatch.
The partial cases are each partial for a reason you can see in the trace. In the program trace a corrupted variable is often overwritten by a later assignment before it matters (0.65). In Dyck a wrong stack is often absorbed by later pops (0.53). And in Tower of Hanoi the model notices:
move 13: [1, 1, 0] -> [[3, 1, 1], [4], [2]] -> Wait, disk 1 cannot be on top of disk 1. Let's re-evaluate move 13.
Current state after move 12: Peg 0: [3, 1], Peg 1: [4], Peg 2: [2].
A corrupted peg configuration makes the next move illegal, and the legality rule is enough for the model to detect the corruption and recompute from the move list in the prompt. Propagation is 0.33.
3SUM is undefined here because the model fails to continue correctly from the gold trace we gave it. Unaided, the models do exhaustive search, while the gold trace tries to point it in the direction of skipping over obviously-bad candidates.
| task, model, depth, k | full prompt → redacted |
|---|---|
| boolean expressionsQwen3.5-9B, d12, k = 6 | 1.00 → 1.00 redacted · no change |
| nested arithmeticQwen3.5-9B, d7, k = 3 | 1.00 → 1.00 redacted · no change |
| S5 composition†DeepSeek-V4-Flash, d16, k = 8 | 0.00 → 0.00 redacted · no change |
| program trace†Qwen3.5-9B, d32, k = 16 | 0.96 → 0.96 redacted · no change |
| blocksworldQwen3.5-9B, d2, k = 1 | 0.94 → 0.76 redacted · −0.18 |
| random lookup table†DeepSeek-V4-Flash, d24, k = 12 | 0.88 → 0.68 redacted · −0.20 |
| CRUXEvalQwen3.5-9B, d6, k = 3 | 0.72 → 0.36 redacted · −0.36 |
| DyckQwen3.5-9B, d8, k = 4 | 0.88 → 0.38 redacted · −0.50 |
| entity tracking* (ours)Qwen3.5-9B, d8, k = 4 | 0.78 → 0.28 redacted · −0.50 |
| cup shufflingQwen3.5-9B, d12, k = 6 | 0.98 → 0.30 redacted · −0.68 |
| tag systemQwen3.5-9B, d2, k = 1 | 0.92 → 0.20 redacted · −0.72 |
On the entity-tracking task, Llama-3-70B cheats by ignoring its trace and recomputing from the prompt. To see if similar shenanigans are occurring for other tasks, we prefill the gold trace through step k and redact the prompt, forcing the model to rely on the reasoning trace. For seven of the tasks, the redaction itself is a confound: with a hole in a list-shaped prompt the model loses the alignment between trace and remaining operators, or stops to comment on the hole, and accuracy falls. We didn't invest in local weights and real prompt blinding, so this is a blind spot in this post.
We could not measure token-level density directly; that needs attention masking. But we estimate that {} and {} are the densest tasks by state-per-token. 3SUM, CRUXEval, entity tracking and blocksworld are too weird to estimate density in this way.
| task | source | state | bits / step | answers | passes 9B · 70B · DS | deepest depth | tok / step | filler rec. | knockout shape | redacted / plain | knockout at d−1 | propagation | self-checking | density |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| nested arithmetic | Suzgun et al. 2022 | values of evaluated sub-expressions | 11.3 | ∞ | ✓ ✓ ✓ | d20 | 86 | +0.00 | decided at end | 1.00 / 1.00 | 0.42 | 0.96 | no | 0.06 |
| cup shuffling | Suzgun et al. 2022 | permutation of objects over people | 2.6 | 3 | ✓ ✓ ✓ | d24 | 27 | -0.08 | decided at end | 0.30 / 0.98 | 0.40 | 0.73 | no | 0.40 |
| program trace† | Nye et al. 2021 | integer variable assignments | 20.0 | ∞ | ✓ ✓ ✓ | d72 | 30 | -0.04 | decided at end | 0.96 / 0.96 | 0.54 (d−2) | 0.65 | no | 0.40 |
| Dyck | Suzgun et al. 2022 | stack contents | 2.0 | ∞ | ✓ ✓ ✓ | d48 | 20 | -0.06 | decided at end | 0.38 / 0.88 | 0.02 | 0.53 | weak | 0.13 |
| tag system | Wu et al. 2025 | tag-system symbol queue | 10.2 | 64 | ✓ ✓ ✓ | d30 | 71 | -0.02 | decided at end | 0.20 / 0.92 | 0.00 | 0.51 | no | 0.10 |
| blocksworld | Stechly et al. 2024 | block stacking configuration | 6.0 | ∞ | ✓ ✓ ✓ | d24 | 206 | +0.02 | early | 0.76 / 0.94 | · | · | n/a | · |
| boolean expressions | Suzgun et al. 2022 | partially reduced boolean expression | 1.0 | 2 | ✓ ✓ ✓ | d12 | 43 | +0.18 | flat | 1.00 / 1.00 | · | · | n/a | · |
| random lookup table† (ergonomic) | Ramesh et al. 2024 | current symbol | 3.3 | 10 | ✓ · ✓ | d48 | 27 | · | decided at end | · | 0.10 | 0.99 | no | 0.15 |
| S5 composition† (ergonomic) | Liu et al. 2022 | permutation of 5 | 6.9 | 120 | ✓ · ✓ | d16 | 62 | · | decided at end | · | 0.04 | 0.89 | no | 0.22 |
| cellular automaton* (ergonomic) | Neary & Woods 2006 | row of 8 binary cells | 8.0 | 256 | ✓ · ✓ | d4 | 371 | · | decided at end | · | 0.00 | 0.88 | no | 0.15 |
| Tower of Hanoi (ergonomic) | Shojaee et al. 2025 | disks on three pegs | 4.8 | ∞ | ✓ · ✓ | d32 | 31 | · | decided at end | · | 0.00 | 0.33 | yes | 0.16 |
| 3SUM† (ergonomic) | Pfau et al. 2024 | enumeration position (implicit) and hit flag | 3.3 | 2 | ✓ · ✓ | d12 | 560 | · | decided at end | · | 0.46 | -0.11 | no | · |
| CRUXEval | Gu et al. 2024 | program variable values | · | ∞ | ✓ ✗ ✓ | d12 | 49 | -0.04 | early | 0.36 / 0.72 | · | · | n/a | · |
| entity tracking* (ours) | Kim & Schuster 2023 | contents of every box | 94.6 | ∞ | ✓ ✗ ✓ | d36 | 23 | -0.04 | unclassified | 0.28 / 0.78 | · | · | n/a | · |
| addition† | Nye et al. 2021 | partial sum digits and carry | 4.3 | ∞ | ✗ ✗ ✓ | d2 | 30 | +0.92 | early | · | · | · | n/a | · |
| multiplication | Dziri et al. 2023 | partial products and running sum | 6.5 | ∞ | ✗ ✗ · | · | · | · | n/a | · | · | · | n/a | · |
Everything above, one row per task in its surviving format. And the recommendation, by situation:
| you want | use | because |
|---|---|---|
| every step load-bearing, no escape hatch | random lookup table (ergonomic), nested arithmetic, S5 composition (ergonomic), cellular automaton | knockout at floor with all but one step; propagation 0.88 to 0.99 |
| the densest tokens | synthetic program trace, cup shuffling | 0.40 of tokens are state; but propagation is partial |
| the cheapest steps | Dyck, random lookup table (ergonomic), entity tracking | under 30 tokens per step on every model |
| very deep chains | synthetic program trace, random lookup table, Dyck, Hanoi (ergonomic), the tag system | pass at depth 30 or more on at least one model |
| to study error detection | Tower of Hanoi (ergonomic) | a corrupted state makes the next move illegal, so the model notices and recomputes from the prompt |
| a prompt you can blind without masking | nested arithmetic, boolean expressions, synthetic program trace | redaction is invisible |
| a bounded state space with a known group structure | S5 composition (ergonomic) | 120 states, NC1-hard, no memorization possible |
| to avoid | addition, multiplication, 3SUM | solved without a scratchpad; enumeration position is unwritten state |
Tasks. Each task is one Python module with the same interface: generate an instance at a depth, solve it, check an answer, and render the published trace format for it. Where a published trace exists it is used byte-for-byte as the few-shot exemplar format; where it does not, the format is mine and the task is starred. Instances are procedurally generated and seeded, so the same instances appear in every condition. Depth is the task's own natural knob (digits, swaps, symbols, executed statements, moves, generations).
Survival rules, fixed before any run: a task passes on a model if at some depth acc(CoT) − acc(no-CoT) ≥ 0.2 with an exact McNemar test on the paired samples at p < 0.05, and acc(CoT) ≥ 0.7 at that same depth. n = 50 per cell, raised to 150 where the 9B was marginal. A task survives the first screen if it passes on at least one model; dagger and starred tasks pass through regardless.
Models. The roster was fixed after the first two runs and before the third: DeepSeek-V4-Flash was added because the 70B underperformed, and it ran on the full grid, not only on the 70B's failures. Both Qwen3.5-9B and DeepSeek-V4-Flash are hybrid thinking models run with thinking disabled per request; the scorer records any hidden reasoning tokens per sample and the count is zero across every run reported here. No-CoT is enforced by a system instruction and a 24-token cap, since chat APIs do not allow a true prefill; leaks are logged and were rare.
Noise. Together's FP8 serving is not bitwise deterministic at temperature 0. Differences of ±0.1 at n = 50 are noise, and I have tried not to read anything into them. No-CoT accuracy is prompt-design-sensitive: on DeepSeek addition, the batch-1 no-CoT prompt scores 0.08 while forcing Answer: after an empty trace scores 1.00. That is why addition's filler "recovery" above is an artifact.
What is not here. Attention-level blinding (mask one step and measure the drop, the "linkage" experiment this whole review was in service of) needs local weights and did not fit the API budget. Batch 3 ran on the 9B only, so the propagation figures have no cross-model check. Boolean expressions, blocksworld, CRUXEval and entity tracking have no step-aligned knockout or propagation rows: batch 3 was run on the tasks whose batch-2 knockout curve was decided-at-the-end, and those four read as answer-available-early (blocksworld, CRUXEval), flat (boolean) or unclassified (entity tracking).
The harness. Inspect AI, one environment variable to change the model, resumable sweeps, cached generations, and a batch-API path at half price. The knockout, filler, redaction and corruption conditions all work by prefilling the assistant turn and letting the model continue, which Together honours. Everything in this post is reproducible from the repository for about the price of a nice dinner.
| task | source | one step is | state carried | depth knob | answer space | format |
|---|---|---|---|---|---|---|
| addition† | Nye et al. 2021 | digit columns | partial sum digits and carry | digits | ∞ | published |
| blocksworld | Stechly et al. 2024 | actions executed | block stacking configuration | blocks | ∞ | published |
| boolean expressions | Suzgun et al. 2022 | named sub-expression resolutions | partially reduced boolean expression | nesting depth | 2 | published |
| cellular automaton* | Neary & Woods 2006 | generations | row of 8 binary cells | depth | 256 | cells rev 3 |
| CRUXEval | Gu et al. 2024 | executed statements | program variable values | executed-line | ∞ | published |
| cup shuffling | Suzgun et al. 2022 | pairwise swaps | permutation of objects over people | n | 3 | published |
| Dyck | Suzgun et al. 2022 | input symbols (stack updates) | stack contents | prefix length, k | ∞ | published |
| entity tracking* | Kim & Schuster 2023 | box operations | contents of every box | numops | ∞ | ours |
| Tower of Hanoi | Shojaee et al. 2025 | moves | disks on three pegs | disks | ∞ | execute |
| multiplication | Dziri et al. 2023 | numbered steps = digits_y*(digits_x+1) | partial products and running sum | digits | ∞ | published |
| nested arithmetic | Suzgun et al. 2022 | sub-expression reductions | values of evaluated sub-expressions | nesting depth | ∞ | published |
| random lookup table† | Ramesh et al. 2024 | table applications | current symbol | function.depth | 10 | ergonomic |
| S5 composition† | Liu et al. 2022 | permutations composed | permutation of 5 | depth | 120 | ergonomic |
| program trace† | Nye et al. 2021 | executed statements | integer variable assignments | depth | ∞ | published |
| 3SUM† | Pfau et al. 2024 | candidate-triple checks in the gold trace | enumeration position (implicit) and hit flag | length | 2 | ergonomic |
| tag system | Wu et al. 2025 | tag-system rewrite steps | tag-system symbol queue | depth | 64 | published |
Claude did lit review, implementation, experiments, tables and charts, and drafted the post. I spent about $35 on inference to run the tasks.