Short answer. A representation can be measurable without being causally important, and an intervention can change behaviour without validating a stable human-readable interpretation of the internal coordinate. Riff Systems treats structural evidence, causal relevance and interpretation as separate evidential claims.
From representation to interpretation
Modern interpretability work often uses probes, directions or activation interventions to identify internal structure and test whether changing activations changes model behaviour. Those methods can provide strong evidence, but different methods answer different questions.
NBS-001 / NBS-002 makes that separation explicit. In the controlled valence study, the structural stage passed after a confounded first experiment was rejected; the frozen direction then produced a selective causal intervention effect; but the prospectively specified fixed-coordinate human interpretation failed on fresh data. The confirmatory endpoint is therefore preserved as structural pass, causal pass, interpretation fail.
Related Riff Systems work
- Non-Biological Semantics (NBS) — foundational methodological paper.
- NBS research programme — programme overview and publication sequence.
- NBS-001 / NBS-002 empirical study — structural, causal and interpretation tests in Pythia-70M.
Questions this work addresses
- Does decodability prove that a model uses a representation?
- Does activation steering demonstrate causal relevance?
- Does causal steering prove that a human-readable concept is the model's mechanism?
- How stable are representation coordinates when context changes?
- What should count as prospective validation of an interpretation?
Context from the wider literature
Activation-steering research treats intervention on internal activations as a way to test and control model behaviour. Recent work also frames steering in explicitly causal terms. NBS uses that broader methodological landscape while imposing an additional separation between successful intervention and transfer of a fixed human interpretation.
External context: Stoehr et al., “Activation Scaling for Steering and Interpreting Language Models” (Findings of EMNLP 2024); Doan et al., “Causal Activation Steering via Sparse Mediation” (Findings of EACL 2026).
