Back to AI Research

AI Research

Reading and Steering Representations of Materials-S... | AI Research

Key Takeaways

  • Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model This research investigates whether large language model...
  • Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics.
  • We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions.
  • In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families.
  • A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison.
Paper AbstractExpand

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
This research investigates whether large language models (LLMs) actually "understand" the physics behind materials science or if they are simply relying on patterns in language. While models like Google’s Gemma-4-E4B-it can provide correct answers to complex scientific questions, the study seeks to determine if these answers are backed by internal representations of physical mechanisms—such as how materials deform, fracture, or react to heat—or if they are merely the result of lexical shortcuts and memorized text.

Decoding Internal Concepts

To see what the model "knows," the researchers used a technique called a "Jacobian lens." This method acts as a measuring tool that looks at the model’s internal hidden states—the high-dimensional data the model processes while thinking—and translates them into a ranking of vocabulary words. By comparing this to a simpler "direct" readout, the researchers found that the model does indeed store scientific concepts internally. However, this knowledge is not uniform; it is highly specific to certain types of materials science problems, such as corrosion or toughness, while performing less effectively on others like fatigue.

Testing Physical Logic

A major challenge in AI interpretability is distinguishing between a model that understands a physical law and one that is just comparing numbers. To test this, the researchers created a "60-law counterfactual benchmark." They presented the model with identical prompts but reversed the physical inputs—for example, asking what happens when a material is refined versus when it is coarsened. They found that the model’s internal state changes were consistent with the governing physical laws. When the researchers reversed the direction of the physical input, the model’s internal representations shifted in a way that tracked the correct constitutive law, proving that the model is tracking the relationship between variables rather than just memorizing static facts.

Steering Model Decisions

Beyond just reading what the model knows, the researchers tested whether they could actively control its output. By performing "causal interventions"—essentially nudging the model’s internal states—they were able to shift the model’s answers toward or away from physically appropriate outcomes. For instance, by patching the model’s internal state with information from a different physical scenario, they could force the model to change its engineering decision. This confirms that specific internal representations are not just passive observations but are actively used by the model to generate its final answers.

Key Takeaways

The study concludes that while materials science information is present and readable within the model, it is not organized in a simple, universal way. Apparent patterns in the model’s "thinking" can sometimes be explained by simple numerical comparisons rather than deep physical insight. Therefore, the researchers emphasize that causal interventions—actually changing the model's internal state to see if the output changes—are essential for verifying that a model is truly using scientific mechanisms to solve problems, rather than relying on shallow associations.

Comments (0)

No comments yet

Be the first to share your thoughts!