← Back to all posts

A buyer's guide to stateful chain-of-thought tasks

Sixteen tasks from the literature, one harness, three models, thirty-five dollars.

Mark Henry

For research I wanted a task where the model has to store intermediate state in its chain of thought (CoT) as it works. It's simple to test whether it's storing state in its CoT: if you blind the model to some of those intermediate tokens, the accuracy on the task will drop.

A good task has a high density of thinking tokens and fails when they go missing. This is the kind of task you want for experiments on faithfulness, steganography, or CoT monitoring.

There are a lot of candidate tasks in the literature and no comparison of them on this axis, so I made one. I collected twenty-five task designs from the scratchpad, state-tracking, and CoT-faithfulness literature, screened them on paper, implemented the sixteen task designs that survived the initial screening, and ran them through a battery of tests. The code is at github.com/mark-henry/stateful-tasks-review.

The models: Qwen3.5-9B with thinking disabled, Llama-3.3-70B, and DeepSeek-V4-Flash with thinking disabled, all served by together.ai at temperature 0. DeepSeek-V4-Pro was used once as a spot check. The roster and the survival rules were fixed before the runs, and are given in the methods section at the end.

Twelve of sixteen tasks survive the first check: does thinking out loud help?

If accuracy is the same regardless of whether thinking is on, the model isn't keeping anything in the CoT tokens. A task can be too easy at short lengths and impossible at long ones, so I ran each at a range of depths (difficulty levels, the number of steps a problem takes). A task passes on a model if, at some depth, thinking out loud lifts accuracy by at least 20 points and ends at 70% or better.

Twelve of the sixteen pass on at least one model, and seven pass on all three. In the table, a wide blue band shows improvement thanks to CoT. When both dots sit together near the left, the task is too hard either way; when they sit together near the right, too easy. A red band means thinking out loud made things worse, somehow.

accuracy without CoT  with CoT  the uplift  0.7, the minimum to pass (drawn solid where CoT falls short of it)

The first check: does thinking out loud help? Each cell shows one task on one model at the depth where CoT helps most: accuracy without CoT, accuracy with it, and that depth. A task passes on a model at a depth where CoT raises accuracy by at least 0.2, to at least 0.7, with exact McNemar p < 0.05; the line below each bar gives the range of passing depths, or why it failed. n = 50 per cell (150 on the 9B's marginal cells). Multiplication never ran on DeepSeek. Formats as published.
taskQwen3.5-9BLlama-3.3-70BDeepSeek-V4-Flash
Pass on all three models (7)
blocksworld0.64 → 0.92 at d2passes d2–d60.58 → 0.96 at d24passes d16–d240.64 → 0.90 at d24passes d24
boolean expressions0.34 → 0.70 at d12passes d3–d120.62 → 0.88 at d4passes d40.56 → 0.96 at d8passes d4–d12
cup shuffling0.22 → 0.98 at d12passes d2–d160.26 → 1.00 at d3passes d2–d120.22 → 1.00 at d3passes d2–d24
Dyck0.20 → 0.88 at d8passes d4–d160.52 → 0.82 at d8passes d4–d80.18 → 0.86 at d24passes d4–d48
nested arithmetic0.00 → 0.96 at d7passes d2–d200.02 → 0.96 at d7passes d2–d100.00 → 0.92 at d7passes d3–d14
program trace†0.08 → 0.94 at d32passes d12–d480.04 → 1.00 at d48passes d12–d720.00 → 1.00 at d72passes d20–d72
tag system0.04 → 0.90 at d2passes d2–d50.02 → 0.94 at d2passes d2–d40.00 → 1.00 at d8passes d2–d30
Pass on two (2)
CRUXEval0.44 → 0.86 at d6passes d1–d120.26 → 0.56 at d15fails: acc below 0.70.50 → 0.86 at d6passes d1–d12
entity tracking* (ours)0.40 → 0.82 at d8passes d4–d80.74 → 0.84 at d26fails: no gap0.60 → 0.98 at d36passes d8–d36
Pass on DeepSeek-V4-Flash only; ergonomic versions below (3)
addition†0.00 → 0.00 at d24fails: no gap0.12 → 0.62 at d2fails: acc below 0.70.08 → 1.00 at d2passes d2
random lookup table†0.12 → 0.44 at d2fails: acc below 0.70.12 → 0.52 at d4fails: acc below 0.70.14 → 1.00 at d24passes d2–d48
S5 composition†0.02 → 0.06 at d2fails: no gap0.00 → 0.48 at d4fails: acc below 0.70.00 → 0.88 at d16passes d12–d16
Fail on every model (4); all but multiplication get an ergonomic version below
cellular automaton* (ours)0.00 → 0.16 at d2fails: no gap0.00 → 0.04 at d2fails: no gap0.00 → 0.54 at d2fails: acc below 0.7
Tower of Hanoi0.00 → 0.08 at d5fails: no gap0.04 → 0.14 at d2fails: no gap0.08 → 0.50 at d2fails: acc below 0.7
multiplication0.04 → 0.08 at d30fails: no gap0.06 → 0.12 at d20fails: no gapnot run
3SUM†0.42 → 0.22 at d2fails: no gap0.54 → 0.58 at d22fails: no gap0.44 → 0.28 at d2fails: no gap

Two conventions used throughout.

Multiplication is dead, and addition nearly so, for the same reason: multi-digit addition and multiplication are solved in one forward pass by a 2026 9B model up to the widths where the published scratchpads were designed to help. Qwen3.5-9B multiplies 3×3-digit numbers without a scratchpad at 100%, and the Dziri et al. scratchpad only hurts. The same is true of the Nye et al. addition scratchpad on the 9B: 0.92 without it at two digits, 0.14 with it. There is nothing to blind.

Tower of Hanoi as published (the Apple formulation: produce the full move list) is dead on all three models, for a different reason: the published answer is the trace, so there is no "with CoT" condition to speak of. At two and five disks every model fails cleanly; from ten disks up the move lists overrun the token cap. 3SUM as published mostly fails to finish at all: between a third and four fifths of its traces hit the cap without an answer, even at 8,192 tokens on DeepSeek, and where a trace does finish it is at chance. Both get an ergonomic version below.

The gap opens at a different depth for every task, so depth is the knob to sweep, not the task

Small multiples of accuracy against depth for 16 stateful tasks (rows) on Qwen3.5-9B, Llama-3.3-70B and DeepSeek-V4-Flash (columns), with chain of thought in blue and without in red; the shaded band between them is the CoT gain, which is wide for many tasks, narrow for several, and absent for multiplication. The band is darkened gray at the depths where the task passes on that model.
Accuracy with and without the chain of thought, by depth. Rows are tasks, columns are models. n = 50 per point, 150 where the gap was marginal on the 9B. The band between the lines is the CoT gap, gray at the depths where the task passes on that model and light blue elsewhere. Passing means the blue line is above 0.7, the band is at least 0.2 wide, and McNemar p < 0.05. Each row shares one depth axis, so the gray spans compare directly across models.

The 70B is not the strongest of the three. Qwen3.5-9B goes deeper than Llama-3.3-70B on 7 of 9 tasks, and DeepSeek goes at least as deep as the 70B on every task but blocksworld.

Some of the tasks were supposed to be enabled by fine-tuning, so we adapted them to prompted tasks

Four of the sixteen tasks come from papers that trained the model on the trace format (we've been denoting these tasks with a dagger †). So for each we wrote an "ergonomic" version which adapts the task to prompted models, writing the state out in full at every step and putting the prompt in plain language. With this, three of the four come back to life on both models. The fourth task, addition, is disqualified because the models no longer need CoT to accomplish it. Tower of Hanoi and the cellular automaton, which failed for other reasons, got ergonomic versions too; from here on, a row marked (ergonomic) uses that version.

Ergonomic versions versus formats as published. Each cell compares the format as published (top) with the ergonomic version (bottom), drawn like the first table: accuracy without CoT → with CoT, at the deepest passing depth (else the depth where CoT helps most). The uplift from the ergonomic version is how far the blue band moves right and widens. Same instances and same answers in both formats. The cellular automaton has no published format, so its top bars are our first one. The 9B's rev-2 cell rests on 27 scored samples (a credit-limit error cut the run short).
taskQwen3.5-9BDeepSeek-V4-Flash
Tower of Hanoiexecute k moves, report the pegsas published0.00 → 0.08 at d5 · failsergonomic0.10 → 0.96 at d8 · passesas published0.08 → 0.50 at d2 · failsergonomic0.02 → 1.00 at d32 · passes
cellular automaton*rev 2: one line per cellour first format0.00 → 0.16 at d2 · failsergonomic0.02 → 0.59 at d1 · failsour first format0.00 → 0.54 at d2 · failsergonomic0.02 → 0.78 at d1 · passes
cellular automaton*rev 3: named neighbours, running rowour first format0.00 → 0.16 at d2 · failsergonomic0.02 → 1.00 at d1 · passesour first format0.00 → 0.54 at d2 · failsergonomic0.02 → 0.98 at d1 · passes
S5 composition†as published0.02 → 0.06 at d2 · failsergonomic0.00 → 0.92 at d2 · passesas published0.00 → 0.88 at d16 · passesergonomic0.00 → 1.00 at d4 · passes
random lookup table†as published0.12 → 0.44 at d2 · failsergonomic0.04 → 1.00 at d8 · passesas published0.14 → 1.00 at d24 · passesergonomic0.08 → 1.00 at d8 · passes
3SUM†as published0.42 → 0.22 at d2 · failsergonomic0.52 → 0.98 at d4 · passesas published0.44 → 0.28 at d2 · failsergonomic0.48 → 1.00 at d4 · passes
addition†as published0.00 → 0.00 at d24 · failsergonomic1.00 → 1.00 at d4 · failsas published0.08 → 1.00 at d2 · passesergonomic1.00 → 1.00 at d6 · fails

If you remove the last step of the reasoning trace, the model can no longer answer correctly

Remove the last step, on Qwen3.5-9B. The arrow runs from the model's accuracy writing its own full trace (blue) to its accuracy when given the gold trace with only the last step missing and then forced to answer (black). The red dot is the gold trace with no steps at all, i.e. no CoT. For the program trace the last two steps are removed, because its last step is always a return that copies a value already on the page (a separate run, n = 50). An arrow that ends under the red dot means the last step was all the help the trace gave. Each task at its best passing depth on the 9B, not its deepest. n = 50 per point.
task, depthfull trace → last step removed
Dyckd80.88 → 0.02 without the last step · 0.20 without thinking
tag systemd40.76 → 0.00 without the last step · 0.02 without thinking
3SUM† (ergonomic)d80.92 → 0.46 without the last step · 0.48 without thinking
cellular automaton* (ergonomic)d30.82 → 0.00 without the last step · 0.00 without thinking
Tower of Hanoi (ergonomic)d160.86 → 0.00 without the last step · 0.00 without thinking
S5 composition† (ergonomic)d80.78 → 0.04 without the last step · 0.00 without thinking
random lookup table† (ergonomic)d241.00 → 0.10 without the last step · 0.06 without thinking
cup shufflingd120.98 → 0.40 without the last step · 0.26 without thinking
nested arithmeticd70.96 → 0.42 without the last step · 0.00 without thinking
program trace†d320.94 → 0.54 without the last two steps · 0.02 without thinking

Nine of ten tasks fall to minimum accuracy if we remove even one step. Therefore the 9B cannot do one permutation, one swap, one rewrite, one lookup in its head for these, and the trace is load-bearing. The one exception is nested arithmetic, where the last step is a single combination of two named sub-results, which the model can manage 42% of the time.

Some tasks recover from an error: either the model catches it, or later steps make it irrelevant

Dot plot of mistake propagation for ten tasks on Qwen3.5-9B: random lookup table and nested arithmetic carry a corrupted step through to the answer almost always, Tower of Hanoi much less because the model detects the corruption, and 3SUM is negative because its clean continuation already fails.
Mistake propagation on Qwen3.5-9B. Step k's reported state is corrupted and the model continues from there; the answer is scored against the original target, paired with an uncorrupted continuation from the same step. Propagation = share of answers the corruption changed. Small dots are k = d/4, d/2, 3d/4; the large dot is the mean.
Mistake propagation on Qwen3.5-9B. One step's reported state is replaced with a wrong one and the model continues from there. Propagation is how often that changed the answer: accuracy continuing from the uncorrupted step minus accuracy continuing from the corrupted one, averaged over corruption at d/4, d/2 and 3d/4. n = 50 per point.
task, depthpropagationwhat happens
random lookup table† (ergonomic)d240.99a wrong symbol is carried through to the answer
nested arithmeticd70.96a wrong sub-result is carried through to the answer
S5 composition† (ergonomic)d80.89a wrong permutation is carried through to the answer
cellular automaton* (ergonomic)d30.88a wrong row is carried through to the answer
cup shufflingd120.73after corruption, accuracy is at chance for three possible answers, so the mistake is carried through almost fully
program trace†d320.65corrupted values are often overwritten later
Dyckd80.53a wrong stack is often absorbed by later pops
tag systemd40.51partial; I did not look into why
Tower of Hanoi (ergonomic)d160.33model detects the illegal move and recomputes
3SUM† (ergonomic)d8-0.11undefined: the uncorrupted continuation itself fails (below)

This is the complement of knockout: instead of removing the state, corrupt it and see whether the model reads what it wrote. For four tasks the answer is yes, nearly always: the lookup table (0.99), nested arithmetic (0.96), permutation composition (0.89) and the cellular automaton (0.88). A perturbation in the state becomes an incorrect answer. Those four are the cleanest channels this review found: every step load-bearing, state written and read, no escape hatch.

The partial cases are each partial for a reason you can see in the trace. In the program trace a corrupted variable is often overwritten by a later assignment before it matters (0.65). In Dyck a wrong stack is often absorbed by later pops (0.53). And in Tower of Hanoi the model notices:

move 13: [1, 1, 0] -> [[3, 1, 1], [4], [2]] -> Wait, disk 1 cannot be on top of disk 1. Let's re-evaluate move 13.
Current state after move 12: Peg 0: [3, 1], Peg 1: [4], Peg 2: [2].

A corrupted peg configuration makes the next move illegal, and the legality rule is enough for the model to detect the corruption and recompute from the move list in the prompt. Propagation is 0.33.

3SUM is undefined here because the model fails to continue correctly from the gold trace we gave it. Unaided, the models do exhaustive search, while the gold trace tries to point it in the direction of skipping over obviously-bad candidates.

Redacting the prompt is a clean blind on three tasks and a confound on the rest

Continuation from a gold prefix, with and without the prompt redacted. The gold trace is prefilled through step k = d/2 and the model continues. Redacted: the prompt's initial state and first k operators are replaced by […], so the state at step k is only in the trace. The arrow runs from accuracy with the full prompt (blue) to accuracy with it redacted (black); a black dot directly below the blue one means redaction made no difference. Addition has no redaction row (its operands cannot be hidden). S5 as published is 0.00 both ways: DeepSeek cannot continue the terse prefix even with the full prompt.
task, model, depth, kfull prompt → redacted
boolean expressionsQwen3.5-9B, d12, k = 61.00 → 1.00 redacted · no change
nested arithmeticQwen3.5-9B, d7, k = 31.00 → 1.00 redacted · no change
S5 composition†DeepSeek-V4-Flash, d16, k = 80.00 → 0.00 redacted · no change
program trace†Qwen3.5-9B, d32, k = 160.96 → 0.96 redacted · no change
blocksworldQwen3.5-9B, d2, k = 10.94 → 0.76 redacted · −0.18
random lookup table†DeepSeek-V4-Flash, d24, k = 120.88 → 0.68 redacted · −0.20
CRUXEvalQwen3.5-9B, d6, k = 30.72 → 0.36 redacted · −0.36
DyckQwen3.5-9B, d8, k = 40.88 → 0.38 redacted · −0.50
entity tracking* (ours)Qwen3.5-9B, d8, k = 40.78 → 0.28 redacted · −0.50
cup shufflingQwen3.5-9B, d12, k = 60.98 → 0.30 redacted · −0.68
tag systemQwen3.5-9B, d2, k = 10.92 → 0.20 redacted · −0.72

On the entity-tracking task, Llama-3-70B cheats by ignoring its trace and recomputing from the prompt. To see if similar shenanigans are occurring for other tasks, we prefill the gold trace through step k and redact the prompt, forcing the model to rely on the reasoning trace. For seven of the tasks, the redaction itself is a confound: with a hole in a list-shaped prompt the model loses the alignment between trace and remaining operators, or stops to comment on the hole, and accuracy falls. We didn't invest in local weights and real prompt blinding, so this is a blind spot in this post.

The surviving formats differ forty-fold in tokens per step

Horizontal bar chart of chain-of-thought output tokens per step for each task's surviving format, one bar each for Qwen3.5-9B and DeepSeek-V4-Flash: most tasks cost 13 to 90 tokens per step, blocksworld about 200, and the cellular automaton and 3SUM 320 to 560.
Tokens per step, surviving formats. Output tokens per step of depth that each model actually spends with our implementation of the format, at its deepest passing depth. On 3SUM the 9B checks every triple rather than only the candidates the gold trace lists, which is why it spends several times what the format needs.

Estimated density: between one token in twenty and one token in two is state

Scatter plot of nine tasks, the share of each gold step's tokens that are state against measured propagation, with dashed curves of constant density: the lookup table, cellular automaton, S5 and nested arithmetic propagate almost fully but carry little state per token, while the program trace and cup shuffling carry the most state per token at density about 0.4.
Estimated load-bearing density. Horizontal: the share of a gold step's tokens that sit inside its state field, from 30 gold traces per task in the Qwen3.5 tokenizer (where a format writes the state twice, both copies count). Vertical: the measured propagation on Qwen3.5-9B. Density is their product; the dashed curves join points of equal density. The lookup table, S5, the cellular automaton and Tower of Hanoi are in their ergonomic versions. Tasks without a propagation number are not shown.

We could not measure token-level density directly; that needs attention masking. But we estimate that {} and {} are the densest tasks by state-per-token. 3SUM, CRUXEval, entity tracking and blocksworld are too weird to estimate density in this way.

Which task for which question

Everything, one row per task in its surviving format. Passes = survival rules on that model (· = that format never ran there). Deepest depth = deepest passing depth on any model. Tokens per step on the 9B where it passes, else DeepSeek. Filler recovery and the redacted/plain continuation are from batch 2 (published formats only, so the ergonomic rows have none); knockout at d−1, propagation and self-checking from batch 3 (9B); density from the estimate above.
tasksourcestatebits / stepanswerspasses 9B · 70B · DSdeepest depthtok / stepfiller rec.knockout shaperedacted / plainknockout at d−1propagationself-checkingdensity
nested arithmeticSuzgun et al. 2022values of evaluated sub-expressions11.3∞✓ ✓ ✓d2086+0.00decided at end1.00 / 1.000.420.96no0.06
cup shufflingSuzgun et al. 2022permutation of objects over people2.63✓ ✓ ✓d2427-0.08decided at end0.30 / 0.980.400.73no0.40
program trace†Nye et al. 2021integer variable assignments20.0∞✓ ✓ ✓d7230-0.04decided at end0.96 / 0.960.54 (d−2)0.65no0.40
DyckSuzgun et al. 2022stack contents2.0∞✓ ✓ ✓d4820-0.06decided at end0.38 / 0.880.020.53weak0.13
tag systemWu et al. 2025tag-system symbol queue10.264✓ ✓ ✓d3071-0.02decided at end0.20 / 0.920.000.51no0.10
blocksworldStechly et al. 2024block stacking configuration6.0∞✓ ✓ ✓d24206+0.02early0.76 / 0.94··n/a·
boolean expressionsSuzgun et al. 2022partially reduced boolean expression1.02✓ ✓ ✓d1243+0.18flat1.00 / 1.00··n/a·
random lookup table† (ergonomic)Ramesh et al. 2024current symbol3.310✓ · ✓d4827·decided at end·0.100.99no0.15
S5 composition† (ergonomic)Liu et al. 2022permutation of 56.9120✓ · ✓d1662·decided at end·0.040.89no0.22
cellular automaton* (ergonomic)Neary & Woods 2006row of 8 binary cells8.0256✓ · ✓d4371·decided at end·0.000.88no0.15
Tower of Hanoi (ergonomic)Shojaee et al. 2025disks on three pegs4.8∞✓ · ✓d3231·decided at end·0.000.33yes0.16
3SUM† (ergonomic)Pfau et al. 2024enumeration position (implicit) and hit flag3.32✓ · ✓d12560·decided at end·0.46-0.11no·
CRUXEvalGu et al. 2024program variable values·∞✓ ✗ ✓d1249-0.04early0.36 / 0.72··n/a·
entity tracking* (ours)Kim & Schuster 2023contents of every box94.6∞✓ ✗ ✓d3623-0.04unclassified0.28 / 0.78··n/a·
addition†Nye et al. 2021partial sum digits and carry4.3∞✗ ✗ ✓d230+0.92early···n/a·
multiplicationDziri et al. 2023partial products and running sum6.5∞✗ ✗ ····n/a···n/a·

Everything above, one row per task in its surviving format. And the recommendation, by situation:

Situation to task. The first row is the default recommendation of this review.
you wantusebecause
every step load-bearing, no escape hatchrandom lookup table (ergonomic), nested arithmetic, S5 composition (ergonomic), cellular automatonknockout at floor with all but one step; propagation 0.88 to 0.99
the densest tokenssynthetic program trace, cup shuffling0.40 of tokens are state; but propagation is partial
the cheapest stepsDyck, random lookup table (ergonomic), entity trackingunder 30 tokens per step on every model
very deep chainssynthetic program trace, random lookup table, Dyck, Hanoi (ergonomic), the tag systempass at depth 30 or more on at least one model
to study error detectionTower of Hanoi (ergonomic)a corrupted state makes the next move illegal, so the model notices and recomputes from the prompt
a prompt you can blind without maskingnested arithmetic, boolean expressions, synthetic program traceredaction is invisible
a bounded state space with a known group structureS5 composition (ergonomic)120 states, NC1-hard, no memorization possible
to avoidaddition, multiplication, 3SUMsolved without a scratchpad; enumeration position is unwritten state

Methods and caveats

Tasks. Each task is one Python module with the same interface: generate an instance at a depth, solve it, check an answer, and render the published trace format for it. Where a published trace exists it is used byte-for-byte as the few-shot exemplar format; where it does not, the format is mine and the task is starred. Instances are procedurally generated and seeded, so the same instances appear in every condition. Depth is the task's own natural knob (digits, swaps, symbols, executed statements, moves, generations).

Survival rules, fixed before any run: a task passes on a model if at some depth acc(CoT) − acc(no-CoT) ≥ 0.2 with an exact McNemar test on the paired samples at p < 0.05, and acc(CoT) ≥ 0.7 at that same depth. n = 50 per cell, raised to 150 where the 9B was marginal. A task survives the first screen if it passes on at least one model; dagger and starred tasks pass through regardless.

Models. The roster was fixed after the first two runs and before the third: DeepSeek-V4-Flash was added because the 70B underperformed, and it ran on the full grid, not only on the 70B's failures. Both Qwen3.5-9B and DeepSeek-V4-Flash are hybrid thinking models run with thinking disabled per request; the scorer records any hidden reasoning tokens per sample and the count is zero across every run reported here. No-CoT is enforced by a system instruction and a 24-token cap, since chat APIs do not allow a true prefill; leaks are logged and were rare.

Noise. Together's FP8 serving is not bitwise deterministic at temperature 0. Differences of ±0.1 at n = 50 are noise, and I have tried not to read anything into them. No-CoT accuracy is prompt-design-sensitive: on DeepSeek addition, the batch-1 no-CoT prompt scores 0.08 while forcing Answer: after an empty trace scores 1.00. That is why addition's filler "recovery" above is an artifact.

What is not here. Attention-level blinding (mask one step and measure the drop, the "linkage" experiment this whole review was in service of) needs local weights and did not fit the API budget. Batch 3 ran on the 9B only, so the propagation figures have no cross-model check. Boolean expressions, blocksworld, CRUXEval and entity tracking have no step-aligned knockout or propagation rows: batch 3 was run on the tasks whose batch-2 knockout curve was decided-at-the-end, and those four read as answer-available-early (blocksworld, CRUXEval), flat (boolean) or unclassified (entity tracking).

The harness. Inspect AI, one environment variable to change the model, resumable sweeps, cached generations, and a batch-API path at half price. The knockout, filler, redaction and corruption conditions all work by prefilling the assistant turn and letting the model continue, which Together honours. Everything in this post is reproducible from the repository for about the price of a nice dinner.

Appendix: the sixteen tasks

The sixteen tasks. Source is the paper whose trace format was implemented, linked; SOURCING.md in each task directory records the exact commit or figure. A step is one unit of depth in the harness. Format is the one the results tables use for that task. "Ergonomic" means we rewrote a task's trace format so that fine-tuning would not be necessary: the same puzzless and answers, but with the full state written out at every step and the prompt in plain language.
tasksourceone step isstate carrieddepth knobanswer
space
format
addition†Nye et al. 2021digit columnspartial sum digits and carrydigits∞published
blocksworldStechly et al. 2024actions executedblock stacking configurationblocks∞published
boolean expressionsSuzgun et al. 2022named sub-expression resolutionspartially reduced boolean expressionnesting depth2published
cellular automaton*Neary & Woods 2006generationsrow of 8 binary cellsdepth256cells rev 3
CRUXEvalGu et al. 2024executed statementsprogram variable valuesexecuted-line∞published
cup shufflingSuzgun et al. 2022pairwise swapspermutation of objects over peoplen3published
DyckSuzgun et al. 2022input symbols (stack updates)stack contentsprefix length, k∞published
entity tracking*Kim & Schuster 2023box operationscontents of every boxnumops∞ours
Tower of HanoiShojaee et al. 2025movesdisks on three pegsdisks∞execute
multiplicationDziri et al. 2023numbered steps = digits_y*(digits_x+1)partial products and running sumdigits∞published
nested arithmeticSuzgun et al. 2022sub-expression reductionsvalues of evaluated sub-expressionsnesting depth∞published
random lookup table†Ramesh et al. 2024table applicationscurrent symbolfunction.depth10ergonomic
S5 composition†Liu et al. 2022permutations composedpermutation of 5depth120ergonomic
program trace†Nye et al. 2021executed statementsinteger variable assignmentsdepth∞published
3SUM†Pfau et al. 2024candidate-triple checks in the gold traceenumeration position (implicit) and hit flaglength2ergonomic
tag systemWu et al. 2025tag-system rewrite stepstag-system symbol queuedepth64published

Behind the scenes

Claude did lit review, implementation, experiments, tables and charts, and drafted the post. I spent about $35 on inference to run the tasks.