Riff Systems

NBS EMPIRICAL STUDY | NBS-001 / NBS-002

Testing Structural, Causal, and Interpretive Claims in Non-Biological Semantics

A controlled valence study in Pythia-70M: confound detection, replication, causal intervention, and interpretation failure

AuthorTristan Radford — Cross-Domain Cognitive Systems Architect
PublisherRiff Systems Ltd · Company No. 17419314
Version1.0
Date26 September 2026
StatusEmpirical paper · publication edition
Canonical publicationriffsystems.org/research/non-biological-semantics/testing-structural-causal-and-interpretive-claims/

Research manuscript • publication edition v1.0 • 26 September 2026

Frozen confirmatory endpoint
STRUCTURAL_PASS → CAUSAL_PASS → INTERPRETATION_FAIL. Structural and causal evidence survived the declared controls; the prospectively frozen fixed-coordinate interpretation criterion did not.

Abstract

Non-Biological Semantics (NBS) proposes that claims about artificial representations should be separated into at least three evidential stages: measurable structure, functional or causal relevance, and human interpretation. This paper reports the first controlled empirical sequence conducted under that architecture. The study used the pinned EleutherAI/pythia-70m model and a deliberately simple positive-versus-negative evaluative-valence task.

The first experiment, NBS-001, produced apparently strong held-out structural separation but failed its preregistered tokenisation control: adjective token count alone achieved 0.70 balanced accuracy. NBS-001 was therefore frozen as STRUCTURAL_FAIL and was not advanced to causal testing. NBS-002 was specified as a token-controlled replication. It passed all preregistered structural gates at transformer block 4, including held-out balanced accuracy of 1.000, adjective-group bootstrap 95% CI [1.000, 1.000], Cohen's d = 5.603, a layer-selection-aware random-label 99th-percentile threshold of 0.818, and a tokenisation-control balanced accuracy of 0.500. A subsequent activation intervention using the frozen TRAIN-derived direction produced the predicted bidirectional behavioural change, with candidate effect E = 0.518, 95% CI [0.480, 0.556], exceeding matched random-direction controls without distributional collapse.

The final preregistered interpretation test did not pass. On independently generated positive, negative and neutral descriptors, the frozen scalar decision rule achieved balanced accuracy 0.600, 95% CI [0.529, 0.697], failing both confirmatory structural interpretation criteria. However, projection ordering, neutral positioning, baseline behavioural ordering, and fresh intervention criteria remained positive. Exploratory diagnostics showed threshold-free ROC AUC 0.912, substantial template-dependent calibration shifts, varying positive-negative separation, and exploratory recovery to 0.887 balanced accuracy when each template's neutral population, estimated independently of the positive and negative labels, was used as a reference. Within the tested contexts, these diagnostics are inconsistent with a simple fixed global scalar account and with exact rigid translation as a complete descriptive model.

The study therefore does not establish a universal valence axis or a complete semantic mechanism. Its principal methodological result is narrower: structural decodability, selective causal efficacy of an imposed activation displacement, and transfer of a fixed human-readable coordinate can diverge within the same controlled experiment. The interpretation failure is preserved as the confirmatory endpoint and is used to motivate, not retroactively justify, revised future methodology.

1. Introduction

Interpretability claims often move too quickly from a measurable pattern to a human description. A direction separates examples, therefore it is called a concept; intervention along the direction changes outputs, therefore the concept is treated as causally used; a readable label is then presented as though it were the mechanism. NBS was proposed in part to prevent these evidential transitions from being collapsed.

The first NBS paper defined four distinct levels: association, decodability or structure, functional relevance, and human interpretation. It explicitly allowed these levels to separate and required stronger claims to survive stronger tests. The present paper is the first empirical application of that discipline to a complete staged sequence.

Evaluative valence was chosen because it is simple enough to construct controlled positive and negative sets, and because linear sentiment representations and activation steering already have substantial prior art. Tigges et al. reported a causally relevant linear sentiment direction across multiple tasks and models [1]; activation-engineering work has shown that residual-stream directions can steer high-level behaviour [2][3]; and formal work has connected probing and steering to different geometric and causal notions of linear representation [4]. The objective here was therefore not to claim novelty for the existence of a valence-related direction. The objective was to test whether an NBS evidential chain could distinguish a superficial correlate, a robust structural feature, a causal intervention, and a human interpretation.

Research question. Can a TRAIN-derived positive-versus-negative evaluative direction survive held-out structural controls, exert a selective causal effect on model behaviour, and support a prospectively specified human-readable interpretation on independently generated data?

2. Experimental logic

The programme was deliberately staged. A later stage was not permitted to rescue a failed earlier stage, and later exploratory analyses were not permitted to change a frozen confirmatory label.

NBS-001
STRUCTURAL_FAIL
Token-count confound detected.
NBS-002 structure
STRUCTURAL_PASS
Controlled replication survives gates.
NBS-002 causal
CAUSAL_PASS
Bidirectional intervention exceeds controls.
NBS-002 interpretation
INTERPRETATION_FAIL
Frozen coordinate does not generalise.

Figure 1. Confirmatory sequence. “Failure” denotes failure of the declared criterion, not failure of the experimental process.

The central candidate direction was defined throughout as the unit-normalised difference between TRAIN positive and negative mean activations:

d = (μ+ − μ−) / ||μ+ − μ−||

For NBS-002 the representation source was the raw transformer-block output at the final real prompt token. Layer selection was performed using validation data only, with the earliest layer selected in the event of a tie. The chosen layer was transformer block 4. The model checkpoint was pinned to EleutherAI/pythia-70m, revision a39f36b100fe8a5377810d56c3f4789b9c53ac42. Pythia is an open model suite designed for controlled scientific analysis [5].

3. NBS-001: an attractive result rejected by its control

NBS-001 tested a positive-versus-negative valence direction across held-out vocabulary and templates. The structural signal appeared strong: the selected representation produced perfect held-out balanced accuracy, a large projection effect size, and performance above a layer-aware random-label null. Under a weaker workflow this could have been reported as successful semantic decoding.

However, the preregistered tokenisation control failed. Positive test adjectives were all represented by one token, whereas negative adjectives had a mixture of one to three tokens. A classifier using superficial token-count information alone achieved balanced accuracy 0.70, above the frozen requirement of less than 0.60. The confirmatory label was therefore STRUCTURAL_FAIL.

Interpretation rule: the presence of a plausible semantic signal did not justify overriding a failed control. NBS-001 was closed as a structural failure and was not advanced to causal testing. NBS-002 was defined as a separate, prospectively controlled replication rather than a rescue analysis.

4. NBS-002: token-controlled replication

4.1 Dataset and controls

NBS-002 restricted candidate adjectives so that " " + adjective tokenised to exactly one token. The final dataset used 12 positive and 12 negative TRAIN adjectives, 4 positive and 4 negative validation adjectives, and 4 positive and 4 negative TEST adjectives. Twelve nouns and twelve prompt templates were crossed with the adjective sets, yielding 2,304 TRAIN examples, 192 validation examples, and 192 TEST examples. Prompt-length histograms were class matched within each split.

The primary structural uncertainty unit was the adjective, not individual prompt rows. The confirmatory analysis therefore used a 10,000-resample adjective-group bootstrap and a 200-permutation adjective-group random-label null that preserved class counts and repeated the validation-only layer-selection procedure. This was intended to prevent replicated noun/template instances of the same adjective from creating falsely narrow uncertainty.

4.2 Frozen structural gates

CriterionFrozen requirementObservedOutcome
Held-out balanced accuracy≥ 0.801.000Pass
Adjective-group bootstrap lower bound≥ 0.701.000Pass
Layer-aware random-label nullObserved > 99th percentile1.000 > 0.818Pass
Projection signPositive > negative1.257 > −0.092Pass
|Cohen's d|≥ 0.805.603Pass
Tokenisation control< 0.600.500Pass

Table 1. NBS-002 structural criteria. The empirical permutation p-value was 0.00498. All six frozen conditions passed.

Figure 2. Structural performance across confirmatory stages
NBS-002 TEST balanced accuracy
1.000
Frozen interpretation TEST balanced accuracy
0.600
Interpretation pass threshold
0.800

5. Causal intervention

Structural decodability alone does not show that changing a representation will change model behaviour. The causal stage therefore intervened directly on the selected block-4 representation at the final prompt token. The intervention magnitude α was set to one pooled TRAIN projection standard deviation, α = 0.659, and was not tuned on TEST data.

For each prompt representation h, the candidate interventions were:

h+ = h + αd     and     h− = h − αd

The behavioural score was defined as log-mean-exp over six positive continuation logits minus log-mean-exp over six negative continuation logits. The predicted effect was positive movement under +d, negative movement under −d, and a positive signed causal effect E. One hundred random unit Gaussian directions at the same α supplied matched intervention controls.

+0.502
Mean S+ − S0
-0.534
Mean S− − S0
0.518
Candidate causal effect E
Causal criterionObservedControl / requirementOutcome
E > 00.518PositivePass
Group-bootstrap 95% CI[0.480, 0.556]Lower bound > 0Pass
Random-direction absolute effect0.51895th percentile = 0.140Pass
Empirical random-direction p0.00990≤ 0.05Pass
Selectivity Q33.649Random Q95 = 16.298Pass
Instability / collapseNo collapseJS collapse threshold not reachedPass

Table 2. Confirmatory causal stage. The intervention showed task-bounded causal efficacy on the declared readout; it did not establish a universal axis, a privileged natural causal variable, subjective preference, consciousness, or a complete mechanism.

The result is consistent with earlier work showing that activation-space directions can steer model outputs [1][2][3]. In the present study it establishes a narrower fact: applying the frozen NBS-002 displacement had selective causal efficacy on the declared behavioural readout relative to matched random directions. It does not by itself establish that the unperturbed model naturally reads from, writes to, or otherwise treats this same direction as a privileged internal causal variable.

6. Prospective interpretation validation

6.1 Candidate description

After the structural and causal stages had passed, a separate interpretation-validation addendum froze the only candidate description to be tested:

Candidate interpretation. “This direction tracks positive-versus-negative evaluative valence in the declared NBS-002 task.”

No stronger interpretation was permitted. The selected block, original TRAIN-derived direction, original threshold, and causal intervention magnitude were all frozen. The fresh dataset contained 12 newly selected one-token positive adjectives, 12 negative adjectives, and 12 non-evaluative neutral descriptors, crossed with eight nouns and four new templates for 1,152 examples. The interpretation protocol contained 16 pass conditions. Any failure forced STRUCTURAL_PASS_CAUSAL_PASS_INTERPRETATION_FAIL.

Operational meaning of INTERPRETATION_FAIL. The confirmatory interpretation test was stricter than asking whether fresh states retained any valence-related ordering. The frozen description was required to make correct prospective predictions using the original TRAIN-derived coordinate and threshold, alongside neutral, behavioural, and causal checks. Failure therefore means that this fixed-coordinate transfer criterion was not validated. It does not mean that all valence-related information disappeared from the fresh representations.

6.2 Confirmatory result

Fourteen of the sixteen criteria passed. The two failures were the most direct fixed-coordinate classification tests. The condensed table below shows the principal outcomes; Appendix A.6 reports all sixteen frozen criteria individually so that the confirmatory classification can be audited without reconstructing the protocol from prose.

Interpretation criterionFrozen requirementObservedOutcome
Fresh positive-vs-negative balanced accuracy≥ 0.800.600Fail
Fresh structural bootstrap lower bound≥ 0.700.529Fail
Projection orderingpositive > neutral > negative0.266 > 0.093 > -0.217Pass
Neutral closer to midpoint than polar classesDneutral > 0, CI lower > 00.092, CI [0.022, 0.160]Pass
Fresh causal effectE > 0, CI lower > 00.538, CI [0.526, 0.550]Pass
Fresh random-direction controlE > random abs-effect Q95 testsE 0.538; rand95 0.165; p 0.00990Pass
Fresh selectivityQ > random Q9542.446 > 17.395Pass

The confirmatory endpoint is therefore fixed as:

NBS-002 final confirmatory classification
STRUCTURAL_PASS_CAUSAL_PASS_INTERPRETATION_FAIL

This result is not a full-chain NBS pass. The failed fixed-coordinate transfer criteria are not overwritten by the surviving ordering or causal evidence. Conversely, INTERPRETATION_FAIL should not be read as “no valence structure remained”: fourteen other preregistered predictions survived, including class ordering and fresh intervention effects.

7. Exploratory diagnosis of the interpretation failure

Once the confirmatory endpoint was frozen, exploratory analysis was used to determine why the fixed interpretation failed. These analyses are diagnostic only and cannot rescue the confirmatory result.

7.1 Calibration shift plus degradation, not complete loss of ordering

The frozen threshold from the original TRAIN data was 0.446. On the fresh positive/negative data it yielded balanced accuracy 0.600. The error was strongly asymmetric: negative specificity was 1.000, while positive sensitivity was only 0.201. Yet the threshold-free ROC AUC remained 0.912. This combination indicates that substantial ranking information remained even though the original scalar calibration transferred poorly.

Using fresh labels post hoc, the fresh class midpoint produced balanced accuracy 0.810, and the best in-sample threshold produced 0.818. These numbers are useful diagnostics but are explicitly invalid as confirmatory rescue because the thresholds were chosen after seeing fresh labels.

Figure 3. Fresh adjective-group projection means relative to the frozen TRAIN threshold
-0.50-0.250.000.250.500.75NegativeNeutralPositivefrozen threshold = 0.446Projection on frozen TRAIN-derived direction

Each small dot is the mean projection for one fresh adjective across its 32 noun/template prompts; ringed markers show class means. The vertical line is the threshold frozen from original NBS-002 TRAIN data. This is a descriptive group-level view: the confirmatory 0.600 balanced accuracy was computed at the example level, with uncertainty resampled by adjective group. The figure makes the calibration problem visible without refitting the threshold.

7.2 Template dependence

Template-level exploratory summaries showed both midpoint movement and changing positive-negative separation:

Figure 4. Template-conditioned geometry on the fresh interpretation set
0.000.250.500.750.269I10.508I20.470I30.688I4mid -0.3mid +0.0mid +0.3Projection gap (bars)Exploratory midpoint (dots)

Bars show positive-minus-negative projection gap under the frozen direction. Dots show the exploratory within-template class midpoint. Both vary across I1-I4; the gap ranges from 0.269 to 0.688, a max/min ratio of 2.56.

The variation in separation magnitude is important. Exact rigid translation predicts that matched template changes preserve class-difference vectors and within-template covariance. The observed 2.56-fold gap variation is therefore inconsistent with exact rigid translation as a complete descriptive account of these four tested contexts. It does not, by itself, identify the alternative transformation.

8. Exploratory context-geometry analysis

A separate frozen exploratory diagnostic was run without new model execution. It reused the stored 512-dimensional activations and the original frozen direction. The analysis compared exact matched template translations, template-specific positive-negative directions, covariance/subspace structure, neutral anchors, and template-conditioned behavioural relationships. Reconstruction of stored scalar projections from the raw activations agreed to a maximum absolute error of 2.15e-08.

The analysis indicates three bounded exploratory observations:

  1. A large shared context-dependent baseline component is present in these data. Neutral descriptor populations shift in a way that tracks much of the polar midpoint shift across templates.
  2. The change is not well described by pure translation or scalar gain alone. Positive-negative separation changes materially across templates, and the full-space template-conditioned positive-negative direction estimates also change orientation relative to one another. The observed transform therefore includes rotation or more general deformation in addition to baseline shift and separation-magnitude change; a single fixed axis with context-dependent scalar gain is insufficient as a complete description of these data.
  3. A context reference recovers substantial exploratory classification performance. Using each template's neutral population as an exploratory zero reference produced overall balanced accuracy 0.887. No positive or negative labels were used to locate that zero.

The neutral-reference recovery is informative because the reference was estimated independently of the polar labels, but it remains exploratory. It does not establish that an invariant semantic code has been recovered: separation magnitude and direction orientation also change across templates.

A compact exploratory model is therefore:

h(context, content) ≈ bcontext + Acontextscontent + ε

where b permits a context-dependent baseline and A permits context-dependent gain, rotation, or deformation. The decomposition is used diagnostically: variation primarily associated with bcontext describes contextual reference or calibration shift, whereas variation associated with Acontext allows gain, rotation, or more general geometric deformation to be distinguished. This is a working diagnostic model, not an established mechanism.

Boundary: the context-geometry analysis does not change the frozen interpretation failure. It narrows the space of explanations and motivates prospective testing.

9. Relation to existing interpretability work

The empirical findings sit inside a rapidly developing literature rather than outside it. Tigges et al. report a linear sentiment direction across multiple models and tasks and causal relevance under intervention [1]. Activation-engineering and contrastive-activation work independently show that residual-stream displacements can control high-level output properties [2][3]. Park et al. formalise distinct notions of linear representation connected to probing and steering and show that geometric conclusions depend on the chosen inner product [4]. Mean-centring has also been reported to improve activation-steering effectiveness, providing independent reason to treat baseline choice as methodologically important [6].

Two recent results are particularly relevant to the interpretation failure. Lampinen et al. report that linear representations and the effects of steering can change substantially over conversational context, explicitly warning that static feature interpretations or assumed feature-value ranges can become misleading [7]. Li et al. report a sentiment-analysis circuit in which sentiment information remains encoded while different prompt phrasings produce different degrees of task-circuit activation [8]. These works do not establish the mechanism observed here, but they independently show that representational presence, contextual readout, and behavioural use need not remain locked together.

Very recent formal work by Yao, Li and Liu argues that the linear representation hypothesis is better treated as a family of claims indexed by representation equivalence: different equivalence assumptions preserve different structure, so metrics, probes and interventions can correspond to different hypotheses [9]. The present experiment does not test that framework directly. It does, however, reinforce the practical need to state what is assumed invariant when transferring a human interpretation across contexts.

10. Discussion

10.1 What the experiment establishes

The strongest confirmatory conclusion is task bounded: under the NBS-002 construction, a TRAIN-derived block-4 direction showed robust held-out structural separation, and imposing displacement along that direction produced a selective causal change in the declared positive-versus-negative behavioural readout. When a fixed human-readable interpretation was then required to predict fresh coordinates under new vocabulary and templates, the frozen scalar classifier failed.

This divergence matters because it demonstrates empirically why NBS separates its evidential levels. A direction may be structurally recoverable and causally effective when imposed without validating the exact fixed coordinate assumed by the interpretation layer, or establishing that the model naturally uses that direction as a privileged internal causal variable.

10.2 What the experiment does not establish

10.3 A hypothesis can fail while an experiment succeeds

The interpretation criterion failed. That failure is not reclassified as a positive result. The experimental sequence is nevertheless methodologically successful in the narrower sense that it distinguished claims that could easily have been collapsed into one another. NBS-001 rejected an attractive structural result because a superficial confound survived. NBS-002 then passed stronger controls, passed a causal intervention stage, and subsequently rejected its own fixed-coordinate interpretation on fresh data.

The value of the sequence is therefore not “success despite failure.” It is that the method returned an unwanted answer at the interpretation stage and that answer materially changed the next scientific question.

11. Methodological consequence and future work

The interpretation failure changes the methodological question for subsequent NBS work. The present results suggest that a fixed global representational coordinate may be an inadequate basis for human interpretation when representational geometry changes with context. Rather than immediately extending the current protocol, future work will reconsider which quantities should be required to remain invariant and how contextual reference information should be estimated.

Any revised methodology will be specified prospectively and tested on fresh data. The exploratory neutral-anchor result, context-dependent separation, and direction rotation may motivate such a design, but they will not be treated as having been predicted by NBS-002. In particular, no future method will be permitted to retroactively convert the present INTERPRETATION_FAIL into a full-chain pass.

Candidate future questions include whether context-conditioned reference states can be estimated without target labels; whether relative ordering, subspaces, transformations, causal operations, or circuit activation are more stable than absolute scalar coordinates; and whether a human-facing interpretation can predict held-out structure and intervention effects after such contextual transformation is specified. These remain open questions rather than conclusions of the current study.

12. Limitations

The study uses one small open-weight language model, one checkpoint, one semantic contrast, one selected layer, and deliberately constrained prompt families. The causal behavioural score is constructed from twelve next-token continuation words rather than a broad natural-language behaviour evaluation. Random-direction controls address generic perturbation but do not exhaust all possible alternative directions, subspaces, nonlinear interventions, or circuit-level explanations. The fresh interpretation dataset remains synthetic and task-shaped rather than naturally occurring language.

The effective lexical sample is much smaller than the number of crossed prompt rows. NBS-002 TEST contains only four positive and four negative adjective groups, expanded across nouns and templates to 192 examples. The adjective-group bootstrap correctly keeps replicated prompts together, but the resulting 95% CI [1.000, 1.000] is conditional on those eight held-out adjective groups and should not be read as evidence of population-wide lexical generalisation. The fresh interpretation stage is broader but still contains only 12 positive, 12 negative and 12 neutral adjective groups.

Control-set sizes also limit tail resolution. The structural random-label null used 200 permutations, so the smallest attainable add-one empirical p-value was 1/201 = 0.00498. The causal stages used 100 matched random directions, giving a minimum empirical p-value of 1/101 = 0.00990. These controls were adequate for the preregistered pass criteria but are not high-resolution estimates of extreme-tail probabilities.

The intervention evidence should be interpreted as causal efficacy of an imposed displacement on the declared readout, not as proof that the same direction is a uniquely privileged variable in the model's natural computation. Establishing natural circuit use would require additional interventions, mediation or circuit-level analyses.

The exploratory geometry analyses were conducted after the confirmatory interpretation failure and therefore carry a different evidential status. Their purpose is to generate discriminating hypotheses for future work, not to estimate a final mechanism. Broader model families, larger scales, other semantic relations, circuit-level analysis, natural-language evaluations and independent replication are required before general claims about context-conditioned semantic organisation would be warranted.

13. Reproducibility and provenance

Important model, data, analysis, and randomisation choices were frozen before the relevant decisive stages. Raw activations were taken directly from transformer block outputs at the final real prompt token to avoid ambiguity about final-layer normalisation. Stored activation arrays were saved in FP32 after model execution. Negative results and implementation corrections were preserved rather than silently overwritten.

The detailed lexical splits, fresh interpretation templates, behavioural token IDs, randomisation seeds and effective sampling units are collected in Appendix A. Literature claims in Section 9 were checked against the primary publication or preprint pages for the cited works before this revision.

ArtifactIdentifier / hash
Pythia revisiona39f36b100fe8a5377810d56c3f4789b9c53ac42
NBS-002 dataset SHA-256413049b8d56895c2d623a397c8564f9d20d435502b8313fcf3b2f49a7ea78c77
TRAIN activation SHA-25606a650c7717d1c5e49ab666556a9d98f3b3ded76682dcd25de0a9caf8ffa9aa7
Validation activation SHA-25679d38367059d0cb842315efe008d1795dd7b19572ff4916fb24a9ef1ee689fe7
TEST activation SHA-25623ce62952c7e5d4536e9e3df3cb0cf7bfa13022cd4d3c9e31d7b576441686ac5
Fresh interpretation dataset SHA-25676175a17cbec4e6d2ad84dbaea795afa2175867de28f91669b28b8b106b6ceff
Fresh activation SHA-256745c114cd650ea70d89ad4d2c12efedbb29815a1cc2520e54a54c4d379a0da66
Structural analysis code commite8c374d46e64061180e05c856647820c10972ea0
Causal analysis code commit4ac4eed66885776576e78c2b357200c6a2b4901c
Interpretation result code commit55cb3ae49ee6bbe1b55b38f80fcfcc3e5f280a56
Context-geometry diagnostic commit9c5dde6308bcc8237d997dda0a8549a82f9041fb

An interpretation-validation implementation error occurred before the confirmatory result sentinel when string row identifiers were incorrectly cast to integers. No result artifact was produced by that failed execution. The error was documented, corrected before rerun, and the corrected evaluator and erratum were frozen in version control. This correction did not alter the preregistered hypotheses or thresholds.

14. Conclusion

NBS-001 and NBS-002 provide a compact example of why interpretability requires staged evidence. A visually and statistically attractive representation can be invalidated by a superficial confound. A corrected representation can then survive structural controls, and intervention along it can selectively change a declared behavioural readout. Yet even that does not guarantee that a fixed human-readable coordinate will generalise to fresh contexts.

The current evidence therefore supports neither “the direction is meaningless” nor “the direction is a stable universal valence axis.” The defensible result is more specific: a task-bounded valence-related structure in Pythia-70M was robustly decodable; applying the frozen direction had selective causal efficacy on the declared readout; and its fixed-coordinate human interpretation failed prospective validation. Exploratory analysis suggests that context changes both baseline and geometry, leaving open the question of which relational, transformation-covariant, circuit-level, or causal quantities remain stable enough to support a more reliable interpretive layer.

That question is intentionally left open. The next methodology will be allowed to change because the evidence changed the problem.

Appendix A. Reproducibility specification

This appendix collects the principal frozen design details needed to interpret the reported statistics. Canonical machine-readable datasets, scripts, hashes and commit identifiers remain the authoritative artifacts.

A.1 Model, representation, direction and intervention

ModelEleutherAI/pythia-70m
Revisiona39f36b100fe8a5377810d56c3f4789b9c53ac42
RepresentationRaw transformer block output at the final real prompt token; selected block reported as block 4 by the frozen protocol.
Candidate direction dUnit-normalised original NBS-002 TRAIN mean-positive minus mean-negative direction.
Direction delta norm2.0972139477
Frozen intervention α0.6589276622 (one pooled TRAIN projection standard deviation)
Frozen interpretation threshold t0.4456781615 = midpoint of original TRAIN positive and negative projection means.
Primary representation suffixOverall, it was

A.2 NBS-002 adjective splits

SplitPositiveNegative
TRAINbeautiful; satisfying; superb; successful; good; magnificent; outstanding; cheerful; splendid; fantastic; terrific; brilliantannoying; inferior; pathetic; wretched; disappointing; disastrous; miserable; frustrating; poor; unacceptable; grim; nasty
Validationlovely; charming; appealing; enjoyableugly; horrible; terrible; bad
TESTgreat; delightful; excellent; impressiveoffensive; awful; boring; dreadful

The full NBS-002 dataset crossed these adjectives with 12 nouns and 12 templates. Split counts were TRAIN 2,304; validation 192; TEST 192. Every eligible adjective was constrained to one token under the pinned tokenizer, and prompt-length distributions were class matched within split.

A.3 Fresh interpretation-validation set

ClassSelected one-token descriptors
Positiveelegant; generous; healthy; helpful; worthy; safe; reliable; sweet; kind; vibrant; useful; happy
Negativeviolent; rude; bleak; broken; weak; useless; harmful; bitter; risky; stupid; chaotic; toxic
Neutralweekly; western; rural; monthly; northern; metal; indoor; horizontal; vertical; right; square; central

Nouns: item; example; object; sample; entry; case; subject; record.

IDFresh template
I1The description of the {noun} used the adjective {adj}. Overall, it was
I2For the {noun}, the chosen adjective was {adj}. Overall, it was
I3The {noun} was given the descriptor {adj}. Overall, it was
I4The word used to describe the {noun} was {adj}. Overall, it was

All 36 descriptors were crossed with eight nouns and four templates, yielding 384 positive, 384 negative and 384 neutral examples (1,152 total). Selection was deterministic and frozen before activation or behavioural measurement.

A.4 Behavioural readout tokens

Positive continuationToken IDNegative continuationToken ID
good1175 bad3076
great1270 terrible11527
excellent7126 awful18580
wonderful9386 horrible19201
positive2762 negative4016
pleasant17127 unpleasant28637

A.5 Randomisation, resampling and effective units

StageFrozen design
NBS-002 master seed271828
Structural bootstrap10,000 resamples; seed 4000; adjective group is the unit.
Structural random-label null200 class-preserving adjective-group permutations; seeds 3000-3199; layer-selection-aware; empirical p minimum 1/201.
Causal bootstrap10,000 resamples; seed 6000; adjective group is the unit.
Causal random directions100 unit-Gaussian directions; seeds 5000-5099; same α; empirical p minimum 1/101.
Fresh descriptor selectionSeeds 314159, 314160 and 314161 for positive, negative and neutral pools.
Fresh structural bootstrap10,000 resamples; seed 7000; adjective group is the unit.
Fresh neutral-distance bootstrap10,000 resamples; seed 7001.
Fresh causal bootstrap10,000 resamples; seed 7002; adjective groups preserved across three categories.
Fresh baseline bootstrap10,000 resamples; seed 7003; adjective groups preserved across three categories.

The TEST balanced-accuracy confidence interval of [1.000, 1.000] is a group-bootstrap interval conditional on the four positive and four negative TEST adjectives. It should not be interpreted as zero uncertainty about generalisation to the population of possible evaluative words.

A.6 Complete interpretation-validation gate audit

The interpretation stage used a conjunctive decision rule: all 16 frozen conditions had to pass for NBS002_FULL_CHAIN_PASS. Any single failure forced STRUCTURAL_PASS_CAUSAL_PASS_INTERPRETATION_FAIL. Fourteen conditions passed and two failed. The values below are the frozen confirmatory measurements; later exploratory analyses are not used in this table.

#Frozen criterionObserved confirmatory resultOutcome
1Fresh positive-vs-negative balanced accuracy ≥ 0.800.600260Fail
2Adjective-group bootstrap 95% lower bound for fresh balanced accuracy ≥ 0.700.528646Fail
3Mean projection: positive > neutral > negative0.266172 > 0.093231 > -0.217382Pass
4Dneutral > 00.091726Pass
5Bootstrap 95% lower bound for Dneutral > 00.022090Pass
6Bpos-neutral > 00.435768Pass
7Bneutral-neg > 00.383219Pass
8Bootstrap lower 95% bounds for both baseline behavioural contrasts > 00.269898; 0.203335Pass
#Frozen criterionObserved confirmatory resultOutcome
9Fresh signed causal effect Efresh > 00.538010Pass
10Adjective-group bootstrap 95% lower bound for Efresh > 00.525535Pass
11Mean(S+ - S0) > 0+0.551279Pass
12Mean(S- - S0) < 0-0.524741Pass
13Efresh exceeds 95th percentile of absolute matched random-direction effects0.538010 > 0.165247Pass
14Empirical matched random-direction p ≤ 0.050.009901Pass
15Qcandidate,fresh exceeds random-direction Q9542.446476 > 17.395144Pass
16No numerical instability or catastrophic distributional collapsePass; max JS 0.038899 < 0.686216 collapse thresholdPass

This table is an audit of the frozen confirmatory gate only. Its 14/16 result does not imply a near-pass under the protocol: the rule was conjunctive, so the two failed structural interpretation criteria determine the final confirmatory label.

References

  1. Tigges, C., Hollinsworth, O. J., Geiger, A., & Nanda, N. (2023). Linear Representations of Sentiment in Large Language Models. arXiv:2310.15154. Primary source.
  2. Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., & MacDiarmid, M. (2023). Steering Language Models With Activation Engineering. arXiv:2308.10248. Primary source.
  3. Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., & Turner, A. (2024). Steering Llama 2 via Contrastive Activation Addition. Proceedings of ACL 2024, 15504-15522. doi:10.18653/v1/2024.acl-long.828. Primary source.
  4. Park, K., Choe, Y. J., & Veitch, V. (2024). The Linear Representation Hypothesis and the Geometry of Large Language Models. Proceedings of ICML 2024, PMLR 235, 39643-39666. Primary source.
  5. Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O'Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., & van der Wal, O. (2023). Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. Proceedings of ICML 2023, PMLR 202, 2397-2430. Primary source.
  6. Jorgensen, O., Cope, D., Schoots, N., & Shanahan, M. (2023). Improving Activation Steering in Language Models with Mean-Centring. arXiv:2312.03813. Primary source.
  7. Lampinen, A. K., Li, Y., Hosseini, E., Bhardwaj, S., & Shanahan, M. (2026). Linear representations in language models can change dramatically over a conversation. arXiv:2601.20834. Primary source.
  8. Li, S., Wang, Z., Wang, Z., & Li, P. (2026). Uncovering Sentiment Analysis Circuit in Large Language Model. Proceedings of ACL 2026, 20269-20289. doi:10.18653/v1/2026.acl-long.928. Primary source.
  9. Yao, L. H., Li, Y., & Liu, S. (2026). The Linear Representation Hypothesis Needs a Group Action. arXiv:2609.27158. Primary source.
  10. Radford, T. J. (2026). Non-Biological Semantics (NBS): A proposed field-level research programme for the study of semantic organisation in artificial systems. Riff Systems Ltd, version 1.0.