← All projects

Hebbian Belief-State World Model

Does a Hebbian synapse state hold readable beliefs about what a model can no longer see?

  • PyTorch
  • World Models
  • Linear probes
  • Mechanistic Interpretation
  • Post Transformer Architecture

software2026

Line chart of probe accuracy against how long an object has been out of view. The LSTM and RWKV lines fall from about 0.69 to the chance line by 65 steps, while the BDH line holds near 0.28.

Problem

Where in an architecture does a belief actually live? A model acting on partial observations has to hold what it can no longer see somewhere in its state, and BDH, the Dragon Hatchling, is a natural place to ask: it carries a fast-changing Hebbian synapse state alongside its slow weights, so the belief has a candidate address rather than being smeared across activations.

Study 1 asked the narrow, falsifiable version: is the belief written there in a linearly readable format? Study 2 asked the fairer one: is it there but written associatively, in a format that a flat probe of 524,288 free parameters cannot estimate from 24,000 examples?

The trap in both is that a large enough state read by a large enough probe decodes something regardless, and a threshold picked after seeing the numbers is always met.

How it was solved

Hypotheses, comparators and numeric thresholds were frozen in a preregistration commit before any run, twice: H1 to H4 for Study 1, H5 to H8 for Study 2, each with a kill criterion that would close the question.

Three parameter-matched models, BDH, LSTM and RWKV at about 1.58 M parameters each, trained to predict observations in a 9 by 9 partially observed gridworld. Probes then read each state and predict the true cell, 81 classes, of an object the agent saw earlier and cannot currently see.

Study 2 added the readouts the architecture’s own addressing implies: query-conditioned rank-r factorizations, derotation before standardizing, and an MLP capacity control. Every family that can be defined on a baseline state was run there too, so nothing was compared against a handicap.

Results

Both kill criteria fired. The synapse state’s best structured readout reaches 0.159 against RWKV’s 0.145 and the LSTM’s 0.113 in the same family, short of the five-point margin over both that the rules required. Chance is 0.011.

Format did matter: Study 2’s H5 passed, worth +0.058 over Study 1’s flat probe. Its own capacity control then undercut the associative reading. An MLP on row norms lands within two points of the structured readout, so the gain attributes to capacity, not to associative structure.

Study 2 also overturned one of Study 1’s verdicts. Belief revision had looked like a failure because the clock started while the object was still visible; measured from the first step it is actually hidden, BDH flips to the new cell within five steps in 77% of episodes and passes.

The model is not broken; the format is the finding, and it is published as a negative result rather than rerun until something passed. Post-hoc, BDH’s probe errors stay spatially local and never worse than chance at any horizon, while both baselines become confidently wrong past 65 steps unseen.