The Reset Button Test
Digital Minds Research Sprint · Track 2

Can an AI outcome signal matter even when it does not last?

We found a task outcome representation in Qwen3 that transfers across several ways of saying success or failure, fades quickly when unrelated context arrives, yet still changes the model's choice to reset a conversation when we intervene on it.

Subramanyam SahooIndependent AI Safety ResearcherAugust 14 to August 16, 2026

The question

AI systems can say they are satisfied, frustrated, confident, or uncertain. Those statements alone do not tell us whether a corresponding internal signal is stable or behaviorally meaningful.

If a model internally represents that a task went well or badly, does that representation generalize, persist, and causally change what the model chooses next?

A deliberately small test

The study uses one open model, Qwen3 4B Instruct 2507, and MMLU questions. The goal is not to prove consciousness. It is to test a measurable chain from task outcome to internal representation to later choice.

1 · Outcome

Pair the same question and reference answer with either a matching or mismatching candidate answer.

2 · Representation

Measure the residual stream difference and select the strongest held out layer.

3 · Transfer

Freeze the direction and test new ways of expressing full credit, scores, reward, and pass or fail.

4 · Decay

Add unrelated tasks and ask how quickly the original outcome information disappears.

5 · Intervention

Steer the frozen direction and measure RESET versus CONTINUE choice probability.

What we found

The central pattern is a separation between natural persistence and causal potency. The signal is easy to read immediately, but it is not a durable internal state.

1.000

Held out AUROC at the selected layer for favorable versus unfavorable task outcomes.

0.998

Mean AUROC at feedback for four previously unseen ways of describing outcome.

0.509

Mean AUROC after eight unrelated task transitions, approximately chance.

+0.025

Reset AUC shift between unfavorable and favorable direction steering in negative outcome contexts.

Reset preference AUC decreases as steering moves toward the favorable outcome direction in both negative and positive outcome contexts.
Causal effect. Moving the residual stream toward the unfavorable outcome direction increases the model's aggregate RESET preference; moving it toward the favorable direction decreases it.
Decodability is initially near one for all feedback families and falls toward chance after unrelated tasks.
Semantic transfer and decay. Credit, numeric score, reward, and verdict wording are almost perfectly decoded at feedback, then quickly lose separability.
Mean retained separation for novel feedback forms falls sharply after one unrelated task and stays near zero.
Not a persistent state. Mean retained geometric separation falls from 1 at feedback to roughly 0.027 after one unrelated task and approximately zero thereafter.
Lexical only outcome phrases retain measurable positive minus negative projection separation across five feedback families.
Words matter, but do not explain everything. Outcome wording alone carries substantial signal. Full task context adds further separation, so the direction is not treated as a privileged welfare variable.
Transient but causally potentThe representation does not behave like a lasting mood. Yet when the same direction is experimentally maintained, it shifts a later choice.

Why this distinction matters

A decodable internal signal should not automatically be treated as a stable welfare state. Conversely, rapid natural decay does not imply the direction is behaviorally irrelevant. The experiment shows why welfare oriented measurement benefits from combining representation analysis, semantic controls, persistence tests, and causal intervention.

What this does not show

The project intentionally keeps its interpretation narrower than the language of suffering or consciousness.

Not evidence of subjective experience

The favorable and unfavorable labels describe externally defined task outcome relations. They are not measurements of phenomenal welfare.

Reset cost is not an economic price

RESET probability was not reliably monotone in the stated percentage cost. We therefore report a cost aggregated preference statistic rather than willingness to pay.

One model and one benchmark

The result is a tightly scoped sprint experiment. Model family replication and stronger environment enacted choices are future work.

Reproduce and inspect

The repository includes the final experiment scripts, earlier pilot versions, result tables, plots, execution logs, environment metadata, and a checksum manifest.

Main experiment

Extract the outcome direction, choose the layer from held out data, run the complete reset cost sweep, and measure vocabulary entropy.

code/current/reset_button_h100_v2.py

Deep persistence follow up

Freeze the direction, test four unseen feedback families, lexical only controls, and forced horizons from zero through eight unrelated tasks.

code/current/deep_persistence_semantic_transfer.py