Mechanistic Interpretability
Gemma 4 knows physics well enough to steer it, and that changes everything
New research on gemma-4-E4B-it reveals that the model's internal representations of materials science mechanisms are causally linked to its answers. The study combines multiple probing and intervention techniques to show that physics knowledge is encoded in state transformations, not just static patterns.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-29 · 3 min read

Large language models can answer scientific questions, but until now it has been hard to tell whether they actually understand the physics or are just mimicking surface patterns. A preprint posted on arXiv on July 22, 2026, tackles this question head-on with the open-weight model google/gemma-4-E4B-it, using a battery of probing and intervention techniques to find out whether materials science mechanisms are present as readable, causally active representations. See also the broader Gemma 4 family launch.
Three forms of representation
The paper argues that mechanic information takes three experimentally separable forms inside the model: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. The researchers combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark, and causal interventions to establish each one.
In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blind identification of 9 of 10 mechanism families. The tenth family was a miss, which the authors attribute to its being defined more by lexical similarity than by a distinct physical concept. This finding aligns with recent work on how models encode structure across scale, as described in the Kimi K3 benchmark analysis.
The graph audit and the direction test

A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. That is a cautionary result: state-space structure can look like physics when it is actually just arithmetic.
To get around that, the team compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations correctly oriented 39 of 40 directional laws, while lexical controls were near chance. Direct, physically neutral and inverse laws across 60 frozen relations were ordered correctly. The upshot: physical relationships are more visible in controlled state changes than in absolute states alone. This mirrors the approach taken in Sakana AI's work on scaling brain-like architectures.
Causal interventions and counterfactuals
Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases. Counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. This is the kind of causal evidence that goes beyond correlation: when the representation changes direction, the answer follows.
The work is relevant beyond materials science. It suggests a general methodology for testing whether a model's physics knowledge is real or shallow, using controlled direction reversals and targeted interventions rather than static probing alone. The emphasis on movement over position is the methodological innovation. This approach builds on earlier attempts to formalize agent knowledge as described in the hierarchical skill-based architecture for AI agents.
What it means for interpretability
Interpretability research has long wrestled with the question of whether LLMs contain genuine scientific understanding or just a good game of pattern matching. This paper lands on the side of "genuine enough to steer." The representations are not merely there in a passive sense, they are causally active and directionally consistent.
The open-weight nature of gemma-4-E4B-it made the experiment possible. Proprietary models rarely let researchers probe and patch individual hidden states at this granularity. That is a structural advantage for the open-source ecosystem: papers like this one are harder to write on closed models. See the broader implications for open-weight model economics.
There is a sobering note in the graph audit result, though. State-space neighborhoods can deceive. A model can organize its representations in a way that mirrors physics without actually using that organization to reason. The direction test found true physical structure underneath, but not every apparent structure survived scrutiny.
The next step, the authors suggest, is to extend the method to other scientific domains and to open-weight models of varying sizes to see whether the same representational patterns hold across scale. Such an extension could benefit from the kind of dynamic context adaptation seen in Jet-Long's bifocal attention approach.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.