Three separate studies — on logical validity, materials mechanisms, and multilingual processing — each break a different link in the chain that runs from “a classifier can read X off the activations” to “the model represents X and I can steer it.” Read together, they turn a binary question (does the model have this feature or not?) into a graded one: which specific rung did the evidence actually pass, and how narrow is the scope of that rung? A probe at 1.000 pairwise accuracy tells you information is accessible. Whether behavior depends on it, whether an intervention along that direction moves anything, and whether the intervention survives new vocabulary are four different experiments.
The three propositions people run together
Suppose you train a logistic regression on the final prompt-token hidden state and it separates valid from invalid arguments almost perfectly. Three claims are now tempting, and they are not the same claim:
1.
Accessibility. The information is linearly readable from that state, under that dataset construction and that readout.
2.
Use. The model’s answer behavior actually depends on that information.
3.
Control. Pushing the state along the probe’s direction reliably changes behavior in the corresponding way, in settings beyond the one you fit it on.
The first does not imply the second, and passing an intervention test in one setting does not settle the third. These are not philosophical cautions; the evidence below measures three specific transitions: validity information stayed decodable on examples a model answered wrong; one probe direction moved verification behavior no more than norm-matched random directions did; and a grain-size direction that reversed correctly under demanding matched controls failed transfer to new material cohorts using different answer words, including a cohort in a reversed physical regime.
Rung one: accessibility, and the behavioral dissociation that stops there
The validity study (When Decodability Is Not Enough) built a synthetic benchmark of matched valid/invalid argument pairs, then fit ℓ2-regularized linear probes to the final prompt-token state at a layer selected on a training partition only. Across five open-weight models and four splits — random, template-held-out, domain-held-out, and inference-family-held-out — matched-pair accuracy hit 1.000. The probe ranked the valid member above the invalid member of every pair, even where global calibration sagged: Mistral-7B managed only 0.771 global AUROC under domain holdout while still getting every pair ordering right.
The interesting move is the correctness-conditioned analysis. Take only the examples the model answered wrong, and ask whether validity is still decodable there. For Llama-3.2, where both gold classes appear in the incorrect subset, AUROC on incorrectly answered examples was 1.000, 0.995, 1.000, and 0.994 across the four splits. The information is sitting in the state while the output goes the other way. That is Behavioral dissociation: accessible validity information does not guarantee a correct answer. Whether the model relies on that information is a separate causal question; an incorrect answer does not settle it. The authors keep the claim narrow, because for models with near-deterministic answer preferences some of these AUROCs are undefined — the incorrect subset contains only one gold class, so there is nothing to rank.
Rung two: what the controls are actually for
Before treating any of this as evidence about a “validity representation,” the same study checked whether lexical or construction regularities could explain the result. Full-prompt TF–IDF reached 0.970 AUROC on the random split and claim-only TF–IDF 0.965 — so near-perfect in-distribution probe accuracy, on its own, does not establish a general validity representation beyond those surface regularities. What distinguishes the hidden-state result is behavior under shift: full-prompt TF–IDF fell to 0.855 under template holdout and 0.814 under domain holdout, claim-only to 0.682 and 0.767, while hidden-state probes generally held up better (Mistral’s 0.771 domain result being the exception). Premises-only classifiers stayed at chance, metadata-only classifiers sat between roughly 0.49 and 0.51, and 200 shuffled-label permutations produced null means of 0.496–0.499, with every model’s observed random-split result exceeding all 200 permutations.
Those controls strengthen the interpretation of what is accessible. Surface regularities contribute substantially to in-distribution separability, but do not fully account for the hidden-state generalization under the tested shifts. Note what each control rules out separately: the lexical baselines address confounds in the data, the shuffled-label permutations address arbitrary linear separability in a high-dimensional space, and the metadata classifier addresses construction artifacts. Drop any one and the remaining evidence supports a weaker sentence.
Rung three: the direction that does almost nothing
Then the study intervened. The normalized probe direction was expressed in raw activation coordinates, scaled by the training-set standard deviation of projection onto it, and added to the final prompt-token state at the selected layer. At α=+4, the VALID–INVALID output margin moved by −0.0037 for Pythia-2.8B, −0.0022 for Llama-3.2, and +0.0023 for Mistral-7B. The signs are inconsistent across models, and five norm-matched random orthogonal directions produced effects of comparable or greater magnitude.
The random-direction comparison is what makes this readable. Without it, a small nonzero margin shift could be reported as “the direction has an effect.” With it, the learned direction shows no consistently signed advantage over the tested norm-matched random orthogonal directions. The authors are careful about what this rules out: it tests whether this particular linear direction, at that site, at that scale, is sufficient to influence verification behavior. Validity information could still be encoded nonlinearly, distributed across token positions, or read at a different computational site, and weak steering does not show that validity-related information is causally irrelevant to the model’s computation.
The constructive contrast: steering that works, then stops working
The validity case tests one probe-derived direction at a selected site in three models. The materials-science study (Reading and Steering Representations of Materials-Science Mechanisms) shows the harder and more useful shape of the problem: an intervention that passes a genuinely demanding causal test and then fails a transfer test.
The setup deserves explanation, because the design is what gives the result force. Hall–Petch relation says that refining grain size increases yield strength, because more grain boundary area impedes dislocation motion. The authors built a direction at layer 16 from earlier data, froze it, and then froze a new six-material matched-pair cohort — titanium, magnesium, low-carbon steel, silver, cobalt alloy, bronze. Each material contributed a refinement prompt and a coarsening prompt with identical material identity, grain sizes, covariates, and answer words; only the direction of change reversed. Titanium went 64→8 micrometers or 8→64, with composition, texture, precipitates, porosity, and dislocation density held fixed.
This asks something stronger than “does the direction favor the word higher.” A context-sensitive grain mechanism must reverse its output effect when only the physical process reverses, with the vector, layer, and answer words unchanged. It did. The raw higher-minus-lower effect was positive after refinement and negative after coarsening in all six pairs; oriented toward the physically correct answer, all 12 conditions were positive and all 6 pairs correct in both relations. The pair-level matched effect was +0.852 log odds (95% interval 0.729–0.988), beating random directions by +0.774, unrelated-mechanism directions by +1.351, and direct controls by +1.043. A post hoc audit of the already-collected dose curves found a +0.484 effect over the smaller ±2% range, positive in all six pairs — though only 10 of 72 stored trajectories were strictly monotone at every adjacent dose, so this supports a consistent local slope rather than a generally linear dose response.
That is a real selective causal result. The authors then asked the next question: does it transfer?
Rung four: transfer to new cohorts fails
Two further six-pair cohorts kept the grain direction unchanged and swapped the answer strings from higher/lower to increase/decrease. In new conventional polycrystals, only 7 of 12 conditions moved correctly and only 1 of 6 pairs was right for both relations. In nanocrystalline materials explicitly stated to lie below the Inverse Hall–Petch regime crossover — where refinement should decrease strength — only 6 of 12 conditions were correct and no pair passed both directions. The registered gates failed.
The diagnostic detail is that across both cohorts, 23 of 24 conditions moved toward the manifest’s designated positive answer word regardless of whether that word was physically appropriate. The original matched reversal rules out a fixed bias toward higher within that cohort. In the transfer cohorts, the same direction instead usually favored the designated positive answer word. Because those cohorts changed materials as well as answer vocabulary, and one also reversed the physical regime, they do not isolate which change caused the failure. The earlier reversal remains convincing inside its frozen format, but the warranted claim narrows to a localized, context-and-vocabulary-dependent pathway. This grain-direction result is bounded by one 4B checkpoint, one late layer, one direction, six matched pairs and the tested transfer cohorts; it does not establish a general control operator.
Do different probes measure the same feature?
All of the above assumes you know what your probe is measuring. The multilingual study (Lingua Franca or Probing Artifact?) shows that even this can be estimator-dependent. Two families of Latent language identification (LLID) method were applied to the same hidden states. Representation-based estimation fits a per-layer Gaussian mixture with one language-conditioned component per candidate language and reads off component posteriors — it asks which language-conditioned region of activation space the state resembles. Decoding-based estimation applies the unembedding matrix to the intermediate state, either raw (Logit lens) or through a per-layer affine map fit to predict the final hidden state (Tuned lens), then converts the vocabulary distribution into language scores.
On the same open-ended prompts, these disagree. The representation probe assigned substantially more mass to the task-relevant language and less to English than raw logit-lens top-p decoding, most visibly in the 50–75% layer window and most pronounced for moderately multilingual models such as Llama-3.1. For the most multilingual models the average probabilities looked closer, but maximum-language estimates still showed a single non-task, non-English language outscoring the task language under the representation probe.
What should a reader take from disagreement? Not that one probe is broken. A hidden state can be geometrically near one language’s cluster without decoding as that language, and can decode into a language whose representation-space signature is not dominant. The study performs no causal interventions and latent language has no ground truth, so the disagreement documents that the probes capture different aspects of multilingual processing — it cannot adjudicate which reflects internal computation. Probe-specific choices (GMM component count, tuned-lens fitting corpus, anchor position, concept-word construction) can shift absolute values; the study also leaves open whether probe disagreement predicts downstream cross-lingual failures. The three studies expose a similar mistake: treating a probe result as stronger evidence than it is. They do not identify a shared feature or internal mechanism.
An evidence ladder for reviewing internal-feature claims
I would use these five checks when reviewing an internal-feature claim:
1.
Specify the measurement. What states, which token position, which layer, which label source, which readout family? The multilingual case shows that a different estimator on identical states can yield a different picture of the feature itself.
2.
Test accessibility against confounds. Held-out template/domain/family splits, lexical baselines, metadata-only classifiers, shuffled-label permutations. A number without these is compatible with surface regularity.
3.
Check decodability against behavior. Does the information survive on examples the model gets wrong? If yes, you have accessibility plus dissociation: the readout does not guarantee correct behavioral alignment. Causal reliance remains a separate question.
4.
Intervene against norm-matched controls. Random orthogonal directions of the same norm are the minimum comparator; unrelated-mechanism and direct controls sharpen it further.
5.
Test transfer. New answer vocabulary, reversed physical or task regime, other layers, other checkpoints, other architectures. This is where the grain-size result changed meaning.
Passing a rung licenses only the claim tested at that rung. Not every study needs every control — but an omitted rung should shrink the sentence in the abstract.
For a monitor, test whether the probe adds useful information when the model answers incorrectly. Then check how that signal holds up under shifts resembling your traffic. A probe that reads a property off activations can be worth deploying as a signal even when steering along its direction does nothing; the validity study’s incorrect-example AUROCs are exactly the case where a monitor adds information the output does not. But if you plan to route, gate, or block on that signal, the shift tests matter more than the headline accuracy, because the lexical baselines degraded under the tested template and domain holdouts while the hidden-state probes largely did not — with Mistral’s 0.771 domain result as an exception, and all of it on a synthetic benchmark. That margin is evidence about shifts resembling the ones tested; it is not a guarantee for your deployment distribution, which you still have to check.
If you are building a steering intervention, a useful follow-up is to keep the vector, layer, material prompts and physical relations fixed while changing only the answer words. That would isolate answer-vocabulary transfer more directly than the reported cohorts, which also changed examples and, in one cohort, the physical regime. Survival would support transfer across those answer words, not a general mechanism. In the reported cohorts, 23 of 24 conditions moved toward the designated positive answer word; this narrows the original result without identifying vocabulary alone as the cause.
And if you are reviewing this literature, the productive question is no longer “did they intervene?” but “which rung, at what scope, against which comparator?” A strong decoder is a good reason to keep auditing. It is not yet a claim about what the model computes, and it is a long way from a control operator you can ship.