Can an AI outcome signal matter even when it does not last?
We found a task outcome representation in Qwen3 that transfers across several ways of saying success or failure, fades quickly when unrelated context arrives, yet still changes the model's choice to reset a conversation when we intervene on it.
The question
AI systems can say they are satisfied, frustrated, confident, or uncertain. Those statements alone do not tell us whether a corresponding internal signal is stable or behaviorally meaningful.
If a model internally represents that a task went well or badly, does that representation generalize, persist, and causally change what the model chooses next?
A deliberately small test
The study uses one open model, Qwen3 4B Instruct 2507, and MMLU questions. The goal is not to prove consciousness. It is to test a measurable chain from task outcome to internal representation to later choice.
Pair the same question and reference answer with either a matching or mismatching candidate answer.
Measure the residual stream difference and select the strongest held out layer.
Freeze the direction and test new ways of expressing full credit, scores, reward, and pass or fail.
Add unrelated tasks and ask how quickly the original outcome information disappears.
Steer the frozen direction and measure RESET versus CONTINUE choice probability.
What we found
The central pattern is a separation between natural persistence and causal potency. The signal is easy to read immediately, but it is not a durable internal state.
Held out AUROC at the selected layer for favorable versus unfavorable task outcomes.
Mean AUROC at feedback for four previously unseen ways of describing outcome.
Mean AUROC after eight unrelated task transitions, approximately chance.
Reset AUC shift between unfavorable and favorable direction steering in negative outcome contexts.




Why this distinction matters
A decodable internal signal should not automatically be treated as a stable welfare state. Conversely, rapid natural decay does not imply the direction is behaviorally irrelevant. The experiment shows why welfare oriented measurement benefits from combining representation analysis, semantic controls, persistence tests, and causal intervention.
What this does not show
The project intentionally keeps its interpretation narrower than the language of suffering or consciousness.
Not evidence of subjective experience
The favorable and unfavorable labels describe externally defined task outcome relations. They are not measurements of phenomenal welfare.
Reset cost is not an economic price
RESET probability was not reliably monotone in the stated percentage cost. We therefore report a cost aggregated preference statistic rather than willingness to pay.
One model and one benchmark
The result is a tightly scoped sprint experiment. Model family replication and stronger environment enacted choices are future work.
Reproduce and inspect
The repository includes the final experiment scripts, earlier pilot versions, result tables, plots, execution logs, environment metadata, and a checksum manifest.
Main experiment
Extract the outcome direction, choose the layer from held out data, run the complete reset cost sweep, and measure vocabulary entropy.
code/current/reset_button_h100_v2.py
Deep persistence follow up
Freeze the direction, test four unseen feedback families, lexical only controls, and forced horizons from zero through eight unrelated tasks.
code/current/deep_persistence_semantic_transfer.py