Abstract
Non-Biological Semantics (NBS) proposes that claims about artificial representations should be separated into at least three evidential stages: measurable structure, functional or causal relevance, and human interpretation. This paper reports the first controlled empirical sequence conducted under that architecture. The study used the pinned EleutherAI/pythia-70m model and a deliberately simple positive-versus-negative evaluative-valence task.
The first experiment, NBS-001, produced apparently strong held-out structural separation but failed its preregistered tokenisation control: adjective token count alone achieved 0.70 balanced accuracy. NBS-001 was therefore frozen as STRUCTURAL_FAIL and was not advanced to causal testing. NBS-002 was specified as a token-controlled replication. It passed all preregistered structural gates at transformer block 4, including held-out balanced accuracy of 1.000, adjective-group bootstrap 95% CI [1.000, 1.000], Cohen's d = 5.603, a layer-selection-aware random-label 99th-percentile threshold of 0.818, and a tokenisation-control balanced accuracy of 0.500. A subsequent activation intervention using the frozen TRAIN-derived direction produced the predicted bidirectional behavioural change, with candidate effect E = 0.518, 95% CI [0.480, 0.556], exceeding matched random-direction controls without distributional collapse.
The final preregistered interpretation test did not pass. On independently generated positive, negative and neutral descriptors, the frozen scalar decision rule achieved balanced accuracy 0.600, 95% CI [0.529, 0.697], failing both confirmatory structural interpretation criteria. However, projection ordering, neutral positioning, baseline behavioural ordering, and fresh intervention criteria remained positive. Exploratory diagnostics showed threshold-free ROC AUC 0.912, substantial template-dependent calibration shifts, varying positive-negative separation, and exploratory recovery to 0.887 balanced accuracy when each template's neutral population, estimated independently of the positive and negative labels, was used as a reference. Within the tested contexts, these diagnostics are inconsistent with a simple fixed global scalar account and with exact rigid translation as a complete descriptive model.
The study therefore does not establish a universal valence axis or a complete semantic mechanism. Its principal methodological result is narrower: structural decodability, selective causal efficacy of an imposed activation displacement, and transfer of a fixed human-readable coordinate can diverge within the same controlled experiment. The interpretation failure is preserved as the confirmatory endpoint and is used to motivate, not retroactively justify, revised future methodology.
1. Introduction
Interpretability claims often move too quickly from a measurable pattern to a human description. A direction separates examples, therefore it is called a concept; intervention along the direction changes outputs, therefore the concept is treated as causally used; a readable label is then presented as though it were the mechanism. NBS was proposed in part to prevent these evidential transitions from being collapsed.
The first NBS paper defined four distinct levels: association, decodability or structure, functional relevance, and human interpretation. It explicitly allowed these levels to separate and required stronger claims to survive stronger tests. The present paper is the first empirical application of that discipline to a complete staged sequence.
Evaluative valence was chosen because it is simple enough to construct controlled positive and negative sets, and because linear sentiment representations and activation steering already have substantial prior art. Tigges et al. reported a causally relevant linear sentiment direction across multiple tasks and models [1]; activation-engineering work has shown that residual-stream directions can steer high-level behaviour [2][3]; and formal work has connected probing and steering to different geometric and causal notions of linear representation [4]. The objective here was therefore not to claim novelty for the existence of a valence-related direction. The objective was to test whether an NBS evidential chain could distinguish a superficial correlate, a robust structural feature, a causal intervention, and a human interpretation.
2. Experimental logic
The programme was deliberately staged. A later stage was not permitted to rescue a failed earlier stage, and later exploratory analyses were not permitted to change a frozen confirmatory label.
Figure 1. Confirmatory sequence. “Failure” denotes failure of the declared criterion, not failure of the experimental process.
The central candidate direction was defined throughout as the unit-normalised difference between TRAIN positive and negative mean activations:
For NBS-002 the representation source was the raw transformer-block output at the final real prompt token. Layer selection was performed using validation data only, with the earliest layer selected in the event of a tie. The chosen layer was transformer block 4. The model checkpoint was pinned to EleutherAI/pythia-70m, revision a39f36b100fe8a5377810d56c3f4789b9c53ac42. Pythia is an open model suite designed for controlled scientific analysis [5].
3. NBS-001: an attractive result rejected by its control
NBS-001 tested a positive-versus-negative valence direction across held-out vocabulary and templates. The structural signal appeared strong: the selected representation produced perfect held-out balanced accuracy, a large projection effect size, and performance above a layer-aware random-label null. Under a weaker workflow this could have been reported as successful semantic decoding.
However, the preregistered tokenisation control failed. Positive test adjectives were all represented by one token, whereas negative adjectives had a mixture of one to three tokens. A classifier using superficial token-count information alone achieved balanced accuracy 0.70, above the frozen requirement of less than 0.60. The confirmatory label was therefore STRUCTURAL_FAIL.
4. NBS-002: token-controlled replication
4.1 Dataset and controls
NBS-002 restricted candidate adjectives so that " " + adjective tokenised to exactly one token. The final dataset used 12 positive and 12 negative TRAIN adjectives, 4 positive and 4 negative validation adjectives, and 4 positive and 4 negative TEST adjectives. Twelve nouns and twelve prompt templates were crossed with the adjective sets, yielding 2,304 TRAIN examples, 192 validation examples, and 192 TEST examples. Prompt-length histograms were class matched within each split.
The primary structural uncertainty unit was the adjective, not individual prompt rows. The confirmatory analysis therefore used a 10,000-resample adjective-group bootstrap and a 200-permutation adjective-group random-label null that preserved class counts and repeated the validation-only layer-selection procedure. This was intended to prevent replicated noun/template instances of the same adjective from creating falsely narrow uncertainty.
4.2 Frozen structural gates
| Criterion | Frozen requirement | Observed | Outcome |
|---|---|---|---|
| Held-out balanced accuracy | ≥ 0.80 | 1.000 | Pass |
| Adjective-group bootstrap lower bound | ≥ 0.70 | 1.000 | Pass |
| Layer-aware random-label null | Observed > 99th percentile | 1.000 > 0.818 | Pass |
| Projection sign | Positive > negative | 1.257 > −0.092 | Pass |
| |Cohen's d| | ≥ 0.80 | 5.603 | Pass |
| Tokenisation control | < 0.60 | 0.500 | Pass |
Table 1. NBS-002 structural criteria. The empirical permutation p-value was 0.00498. All six frozen conditions passed.
5. Causal intervention
Structural decodability alone does not show that changing a representation will change model behaviour. The causal stage therefore intervened directly on the selected block-4 representation at the final prompt token. The intervention magnitude α was set to one pooled TRAIN projection standard deviation, α = 0.659, and was not tuned on TEST data.
For each prompt representation h, the candidate interventions were:
The behavioural score was defined as log-mean-exp over six positive continuation logits minus log-mean-exp over six negative continuation logits. The predicted effect was positive movement under +d, negative movement under −d, and a positive signed causal effect E. One hundred random unit Gaussian directions at the same α supplied matched intervention controls.
| Causal criterion | Observed | Control / requirement | Outcome |
|---|---|---|---|
| E > 0 | 0.518 | Positive | Pass |
| Group-bootstrap 95% CI | [0.480, 0.556] | Lower bound > 0 | Pass |
| Random-direction absolute effect | 0.518 | 95th percentile = 0.140 | Pass |
| Empirical random-direction p | 0.00990 | ≤ 0.05 | Pass |
| Selectivity Q | 33.649 | Random Q95 = 16.298 | Pass |
| Instability / collapse | No collapse | JS collapse threshold not reached | Pass |
Table 2. Confirmatory causal stage. The intervention showed task-bounded causal efficacy on the declared readout; it did not establish a universal axis, a privileged natural causal variable, subjective preference, consciousness, or a complete mechanism.
The result is consistent with earlier work showing that activation-space directions can steer model outputs [1][2][3]. In the present study it establishes a narrower fact: applying the frozen NBS-002 displacement had selective causal efficacy on the declared behavioural readout relative to matched random directions. It does not by itself establish that the unperturbed model naturally reads from, writes to, or otherwise treats this same direction as a privileged internal causal variable.
6. Prospective interpretation validation
6.1 Candidate description
After the structural and causal stages had passed, a separate interpretation-validation addendum froze the only candidate description to be tested:
No stronger interpretation was permitted. The selected block, original TRAIN-derived direction, original threshold, and causal intervention magnitude were all frozen. The fresh dataset contained 12 newly selected one-token positive adjectives, 12 negative adjectives, and 12 non-evaluative neutral descriptors, crossed with eight nouns and four new templates for 1,152 examples. The interpretation protocol contained 16 pass conditions. Any failure forced STRUCTURAL_PASS_CAUSAL_PASS_INTERPRETATION_FAIL.
6.2 Confirmatory result
Fourteen of the sixteen criteria passed. The two failures were the most direct fixed-coordinate classification tests. The condensed table below shows the principal outcomes; Appendix A.6 reports all sixteen frozen criteria individually so that the confirmatory classification can be audited without reconstructing the protocol from prose.
| Interpretation criterion | Frozen requirement | Observed | Outcome |
|---|---|---|---|
| Fresh positive-vs-negative balanced accuracy | ≥ 0.80 | 0.600 | Fail |
| Fresh structural bootstrap lower bound | ≥ 0.70 | 0.529 | Fail |
| Projection ordering | positive > neutral > negative | 0.266 > 0.093 > -0.217 | Pass |
| Neutral closer to midpoint than polar classes | Dneutral > 0, CI lower > 0 | 0.092, CI [0.022, 0.160] | Pass |
| Fresh causal effect | E > 0, CI lower > 0 | 0.538, CI [0.526, 0.550] | Pass |
| Fresh random-direction control | E > random abs-effect Q95 tests | E 0.538; rand95 0.165; p 0.00990 | Pass |
| Fresh selectivity | Q > random Q95 | 42.446 > 17.395 | Pass |
The confirmatory endpoint is therefore fixed as:
This result is not a full-chain NBS pass. The failed fixed-coordinate transfer criteria are not overwritten by the surviving ordering or causal evidence. Conversely, INTERPRETATION_FAIL should not be read as “no valence structure remained”: fourteen other preregistered predictions survived, including class ordering and fresh intervention effects.
7. Exploratory diagnosis of the interpretation failure
Once the confirmatory endpoint was frozen, exploratory analysis was used to determine why the fixed interpretation failed. These analyses are diagnostic only and cannot rescue the confirmatory result.
7.1 Calibration shift plus degradation, not complete loss of ordering
The frozen threshold from the original TRAIN data was 0.446. On the fresh positive/negative data it yielded balanced accuracy 0.600. The error was strongly asymmetric: negative specificity was 1.000, while positive sensitivity was only 0.201. Yet the threshold-free ROC AUC remained 0.912. This combination indicates that substantial ranking information remained even though the original scalar calibration transferred poorly.
Using fresh labels post hoc, the fresh class midpoint produced balanced accuracy 0.810, and the best in-sample threshold produced 0.818. These numbers are useful diagnostics but are explicitly invalid as confirmatory rescue because the thresholds were chosen after seeing fresh labels.
Each small dot is the mean projection for one fresh adjective across its 32 noun/template prompts; ringed markers show class means. The vertical line is the threshold frozen from original NBS-002 TRAIN data. This is a descriptive group-level view: the confirmatory 0.600 balanced accuracy was computed at the example level, with uncertainty resampled by adjective group. The figure makes the calibration problem visible without refitting the threshold.
7.2 Template dependence
Template-level exploratory summaries showed both midpoint movement and changing positive-negative separation:
Bars show positive-minus-negative projection gap under the frozen direction. Dots show the exploratory within-template class midpoint. Both vary across I1-I4; the gap ranges from 0.269 to 0.688, a max/min ratio of 2.56.
The variation in separation magnitude is important. Exact rigid translation predicts that matched template changes preserve class-difference vectors and within-template covariance. The observed 2.56-fold gap variation is therefore inconsistent with exact rigid translation as a complete descriptive account of these four tested contexts. It does not, by itself, identify the alternative transformation.
8. Exploratory context-geometry analysis
A separate frozen exploratory diagnostic was run without new model execution. It reused the stored 512-dimensional activations and the original frozen direction. The analysis compared exact matched template translations, template-specific positive-negative directions, covariance/subspace structure, neutral anchors, and template-conditioned behavioural relationships. Reconstruction of stored scalar projections from the raw activations agreed to a maximum absolute error of 2.15e-08.
The analysis indicates three bounded exploratory observations:
- A large shared context-dependent baseline component is present in these data. Neutral descriptor populations shift in a way that tracks much of the polar midpoint shift across templates.
- The change is not well described by pure translation or scalar gain alone. Positive-negative separation changes materially across templates, and the full-space template-conditioned positive-negative direction estimates also change orientation relative to one another. The observed transform therefore includes rotation or more general deformation in addition to baseline shift and separation-magnitude change; a single fixed axis with context-dependent scalar gain is insufficient as a complete description of these data.
- A context reference recovers substantial exploratory classification performance. Using each template's neutral population as an exploratory zero reference produced overall balanced accuracy 0.887. No positive or negative labels were used to locate that zero.
The neutral-reference recovery is informative because the reference was estimated independently of the polar labels, but it remains exploratory. It does not establish that an invariant semantic code has been recovered: separation magnitude and direction orientation also change across templates.
A compact exploratory model is therefore:
where b permits a context-dependent baseline and A permits context-dependent gain, rotation, or deformation. The decomposition is used diagnostically: variation primarily associated with bcontext describes contextual reference or calibration shift, whereas variation associated with Acontext allows gain, rotation, or more general geometric deformation to be distinguished. This is a working diagnostic model, not an established mechanism.
9. Relation to existing interpretability work
The empirical findings sit inside a rapidly developing literature rather than outside it. Tigges et al. report a linear sentiment direction across multiple models and tasks and causal relevance under intervention [1]. Activation-engineering and contrastive-activation work independently show that residual-stream displacements can control high-level output properties [2][3]. Park et al. formalise distinct notions of linear representation connected to probing and steering and show that geometric conclusions depend on the chosen inner product [4]. Mean-centring has also been reported to improve activation-steering effectiveness, providing independent reason to treat baseline choice as methodologically important [6].
Two recent results are particularly relevant to the interpretation failure. Lampinen et al. report that linear representations and the effects of steering can change substantially over conversational context, explicitly warning that static feature interpretations or assumed feature-value ranges can become misleading [7]. Li et al. report a sentiment-analysis circuit in which sentiment information remains encoded while different prompt phrasings produce different degrees of task-circuit activation [8]. These works do not establish the mechanism observed here, but they independently show that representational presence, contextual readout, and behavioural use need not remain locked together.
Very recent formal work by Yao, Li and Liu argues that the linear representation hypothesis is better treated as a family of claims indexed by representation equivalence: different equivalence assumptions preserve different structure, so metrics, probes and interventions can correspond to different hypotheses [9]. The present experiment does not test that framework directly. It does, however, reinforce the practical need to state what is assumed invariant when transferring a human interpretation across contexts.
10. Discussion
10.1 What the experiment establishes
The strongest confirmatory conclusion is task bounded: under the NBS-002 construction, a TRAIN-derived block-4 direction showed robust held-out structural separation, and imposing displacement along that direction produced a selective causal change in the declared positive-versus-negative behavioural readout. When a fixed human-readable interpretation was then required to predict fresh coordinates under new vocabulary and templates, the frozen scalar classifier failed.
This divergence matters because it demonstrates empirically why NBS separates its evidential levels. A direction may be structurally recoverable and causally effective when imposed without validating the exact fixed coordinate assumed by the interpretation layer, or establishing that the model naturally uses that direction as a privileged internal causal variable.
10.2 What the experiment does not establish
- It does not establish a universal valence representation.
- It does not establish the model's exhaustive or “true” concept of valence.
- It does not establish subjective meaning, preference, emotion, sentience, or consciousness.
- It does not establish that the exploratory context transform is the actual model mechanism.
- It does not establish novelty for linear sentiment representation or activation steering.
- It does not validate NBS as an independent scientific field.
- It does not constitute a human-subject validation of an NBS-II interface.
10.3 A hypothesis can fail while an experiment succeeds
The interpretation criterion failed. That failure is not reclassified as a positive result. The experimental sequence is nevertheless methodologically successful in the narrower sense that it distinguished claims that could easily have been collapsed into one another. NBS-001 rejected an attractive structural result because a superficial confound survived. NBS-002 then passed stronger controls, passed a causal intervention stage, and subsequently rejected its own fixed-coordinate interpretation on fresh data.
The value of the sequence is therefore not “success despite failure.” It is that the method returned an unwanted answer at the interpretation stage and that answer materially changed the next scientific question.
11. Methodological consequence and future work
The interpretation failure changes the methodological question for subsequent NBS work. The present results suggest that a fixed global representational coordinate may be an inadequate basis for human interpretation when representational geometry changes with context. Rather than immediately extending the current protocol, future work will reconsider which quantities should be required to remain invariant and how contextual reference information should be estimated.
Any revised methodology will be specified prospectively and tested on fresh data. The exploratory neutral-anchor result, context-dependent separation, and direction rotation may motivate such a design, but they will not be treated as having been predicted by NBS-002. In particular, no future method will be permitted to retroactively convert the present INTERPRETATION_FAIL into a full-chain pass.
Candidate future questions include whether context-conditioned reference states can be estimated without target labels; whether relative ordering, subspaces, transformations, causal operations, or circuit activation are more stable than absolute scalar coordinates; and whether a human-facing interpretation can predict held-out structure and intervention effects after such contextual transformation is specified. These remain open questions rather than conclusions of the current study.
12. Limitations
The study uses one small open-weight language model, one checkpoint, one semantic contrast, one selected layer, and deliberately constrained prompt families. The causal behavioural score is constructed from twelve next-token continuation words rather than a broad natural-language behaviour evaluation. Random-direction controls address generic perturbation but do not exhaust all possible alternative directions, subspaces, nonlinear interventions, or circuit-level explanations. The fresh interpretation dataset remains synthetic and task-shaped rather than naturally occurring language.
The effective lexical sample is much smaller than the number of crossed prompt rows. NBS-002 TEST contains only four positive and four negative adjective groups, expanded across nouns and templates to 192 examples. The adjective-group bootstrap correctly keeps replicated prompts together, but the resulting 95% CI [1.000, 1.000] is conditional on those eight held-out adjective groups and should not be read as evidence of population-wide lexical generalisation. The fresh interpretation stage is broader but still contains only 12 positive, 12 negative and 12 neutral adjective groups.
Control-set sizes also limit tail resolution. The structural random-label null used 200 permutations, so the smallest attainable add-one empirical p-value was 1/201 = 0.00498. The causal stages used 100 matched random directions, giving a minimum empirical p-value of 1/101 = 0.00990. These controls were adequate for the preregistered pass criteria but are not high-resolution estimates of extreme-tail probabilities.
The intervention evidence should be interpreted as causal efficacy of an imposed displacement on the declared readout, not as proof that the same direction is a uniquely privileged variable in the model's natural computation. Establishing natural circuit use would require additional interventions, mediation or circuit-level analyses.
The exploratory geometry analyses were conducted after the confirmatory interpretation failure and therefore carry a different evidential status. Their purpose is to generate discriminating hypotheses for future work, not to estimate a final mechanism. Broader model families, larger scales, other semantic relations, circuit-level analysis, natural-language evaluations and independent replication are required before general claims about context-conditioned semantic organisation would be warranted.
13. Reproducibility and provenance
Important model, data, analysis, and randomisation choices were frozen before the relevant decisive stages. Raw activations were taken directly from transformer block outputs at the final real prompt token to avoid ambiguity about final-layer normalisation. Stored activation arrays were saved in FP32 after model execution. Negative results and implementation corrections were preserved rather than silently overwritten.
The detailed lexical splits, fresh interpretation templates, behavioural token IDs, randomisation seeds and effective sampling units are collected in Appendix A. Literature claims in Section 9 were checked against the primary publication or preprint pages for the cited works before this revision.
| Artifact | Identifier / hash |
|---|---|
| Pythia revision | a39f36b100fe8a5377810d56c3f4789b9c53ac42 |
| NBS-002 dataset SHA-256 | 413049b8d56895c2d623a397c8564f9d20d435502b8313fcf3b2f49a7ea78c77 |
| TRAIN activation SHA-256 | 06a650c7717d1c5e49ab666556a9d98f3b3ded76682dcd25de0a9caf8ffa9aa7 |
| Validation activation SHA-256 | 79d38367059d0cb842315efe008d1795dd7b19572ff4916fb24a9ef1ee689fe7 |
| TEST activation SHA-256 | 23ce62952c7e5d4536e9e3df3cb0cf7bfa13022cd4d3c9e31d7b576441686ac5 |
| Fresh interpretation dataset SHA-256 | 76175a17cbec4e6d2ad84dbaea795afa2175867de28f91669b28b8b106b6ceff |
| Fresh activation SHA-256 | 745c114cd650ea70d89ad4d2c12efedbb29815a1cc2520e54a54c4d379a0da66 |
| Structural analysis code commit | e8c374d46e64061180e05c856647820c10972ea0 |
| Causal analysis code commit | 4ac4eed66885776576e78c2b357200c6a2b4901c |
| Interpretation result code commit | 55cb3ae49ee6bbe1b55b38f80fcfcc3e5f280a56 |
| Context-geometry diagnostic commit | 9c5dde6308bcc8237d997dda0a8549a82f9041fb |
An interpretation-validation implementation error occurred before the confirmatory result sentinel when string row identifiers were incorrectly cast to integers. No result artifact was produced by that failed execution. The error was documented, corrected before rerun, and the corrected evaluator and erratum were frozen in version control. This correction did not alter the preregistered hypotheses or thresholds.
14. Conclusion
NBS-001 and NBS-002 provide a compact example of why interpretability requires staged evidence. A visually and statistically attractive representation can be invalidated by a superficial confound. A corrected representation can then survive structural controls, and intervention along it can selectively change a declared behavioural readout. Yet even that does not guarantee that a fixed human-readable coordinate will generalise to fresh contexts.
The current evidence therefore supports neither “the direction is meaningless” nor “the direction is a stable universal valence axis.” The defensible result is more specific: a task-bounded valence-related structure in Pythia-70M was robustly decodable; applying the frozen direction had selective causal efficacy on the declared readout; and its fixed-coordinate human interpretation failed prospective validation. Exploratory analysis suggests that context changes both baseline and geometry, leaving open the question of which relational, transformation-covariant, circuit-level, or causal quantities remain stable enough to support a more reliable interpretive layer.
That question is intentionally left open. The next methodology will be allowed to change because the evidence changed the problem.
Appendix A. Reproducibility specification
This appendix collects the principal frozen design details needed to interpret the reported statistics. Canonical machine-readable datasets, scripts, hashes and commit identifiers remain the authoritative artifacts.
A.1 Model, representation, direction and intervention
| Model | EleutherAI/pythia-70m |
|---|---|
| Revision | a39f36b100fe8a5377810d56c3f4789b9c53ac42 |
| Representation | Raw transformer block output at the final real prompt token; selected block reported as block 4 by the frozen protocol. |
| Candidate direction d | Unit-normalised original NBS-002 TRAIN mean-positive minus mean-negative direction. |
| Direction delta norm | 2.0972139477 |
| Frozen intervention α | 0.6589276622 (one pooled TRAIN projection standard deviation) |
| Frozen interpretation threshold t | 0.4456781615 = midpoint of original TRAIN positive and negative projection means. |
| Primary representation suffix | Overall, it was |
A.2 NBS-002 adjective splits
| Split | Positive | Negative |
|---|---|---|
| TRAIN | beautiful; satisfying; superb; successful; good; magnificent; outstanding; cheerful; splendid; fantastic; terrific; brilliant | annoying; inferior; pathetic; wretched; disappointing; disastrous; miserable; frustrating; poor; unacceptable; grim; nasty |
| Validation | lovely; charming; appealing; enjoyable | ugly; horrible; terrible; bad |
| TEST | great; delightful; excellent; impressive | offensive; awful; boring; dreadful |
The full NBS-002 dataset crossed these adjectives with 12 nouns and 12 templates. Split counts were TRAIN 2,304; validation 192; TEST 192. Every eligible adjective was constrained to one token under the pinned tokenizer, and prompt-length distributions were class matched within split.
A.3 Fresh interpretation-validation set
| Class | Selected one-token descriptors |
|---|---|
| Positive | elegant; generous; healthy; helpful; worthy; safe; reliable; sweet; kind; vibrant; useful; happy |
| Negative | violent; rude; bleak; broken; weak; useless; harmful; bitter; risky; stupid; chaotic; toxic |
| Neutral | weekly; western; rural; monthly; northern; metal; indoor; horizontal; vertical; right; square; central |
Nouns: item; example; object; sample; entry; case; subject; record.
| ID | Fresh template |
|---|---|
| I1 | The description of the {noun} used the adjective {adj}. Overall, it was |
| I2 | For the {noun}, the chosen adjective was {adj}. Overall, it was |
| I3 | The {noun} was given the descriptor {adj}. Overall, it was |
| I4 | The word used to describe the {noun} was {adj}. Overall, it was |
All 36 descriptors were crossed with eight nouns and four templates, yielding 384 positive, 384 negative and 384 neutral examples (1,152 total). Selection was deterministic and frozen before activation or behavioural measurement.
A.4 Behavioural readout tokens
| Positive continuation | Token ID | Negative continuation | Token ID |
|---|---|---|---|
good | 1175 | bad | 3076 |
great | 1270 | terrible | 11527 |
excellent | 7126 | awful | 18580 |
wonderful | 9386 | horrible | 19201 |
positive | 2762 | negative | 4016 |
pleasant | 17127 | unpleasant | 28637 |
A.5 Randomisation, resampling and effective units
| Stage | Frozen design |
|---|---|
| NBS-002 master seed | 271828 |
| Structural bootstrap | 10,000 resamples; seed 4000; adjective group is the unit. |
| Structural random-label null | 200 class-preserving adjective-group permutations; seeds 3000-3199; layer-selection-aware; empirical p minimum 1/201. |
| Causal bootstrap | 10,000 resamples; seed 6000; adjective group is the unit. |
| Causal random directions | 100 unit-Gaussian directions; seeds 5000-5099; same α; empirical p minimum 1/101. |
| Fresh descriptor selection | Seeds 314159, 314160 and 314161 for positive, negative and neutral pools. |
| Fresh structural bootstrap | 10,000 resamples; seed 7000; adjective group is the unit. |
| Fresh neutral-distance bootstrap | 10,000 resamples; seed 7001. |
| Fresh causal bootstrap | 10,000 resamples; seed 7002; adjective groups preserved across three categories. |
| Fresh baseline bootstrap | 10,000 resamples; seed 7003; adjective groups preserved across three categories. |
The TEST balanced-accuracy confidence interval of [1.000, 1.000] is a group-bootstrap interval conditional on the four positive and four negative TEST adjectives. It should not be interpreted as zero uncertainty about generalisation to the population of possible evaluative words.
A.6 Complete interpretation-validation gate audit
The interpretation stage used a conjunctive decision rule: all 16 frozen conditions had to pass for NBS002_FULL_CHAIN_PASS. Any single failure forced STRUCTURAL_PASS_CAUSAL_PASS_INTERPRETATION_FAIL. Fourteen conditions passed and two failed. The values below are the frozen confirmatory measurements; later exploratory analyses are not used in this table.
| # | Frozen criterion | Observed confirmatory result | Outcome |
|---|---|---|---|
| 1 | Fresh positive-vs-negative balanced accuracy ≥ 0.80 | 0.600260 | Fail |
| 2 | Adjective-group bootstrap 95% lower bound for fresh balanced accuracy ≥ 0.70 | 0.528646 | Fail |
| 3 | Mean projection: positive > neutral > negative | 0.266172 > 0.093231 > -0.217382 | Pass |
| 4 | Dneutral > 0 | 0.091726 | Pass |
| 5 | Bootstrap 95% lower bound for Dneutral > 0 | 0.022090 | Pass |
| 6 | Bpos-neutral > 0 | 0.435768 | Pass |
| 7 | Bneutral-neg > 0 | 0.383219 | Pass |
| 8 | Bootstrap lower 95% bounds for both baseline behavioural contrasts > 0 | 0.269898; 0.203335 | Pass |
| # | Frozen criterion | Observed confirmatory result | Outcome |
|---|---|---|---|
| 9 | Fresh signed causal effect Efresh > 0 | 0.538010 | Pass |
| 10 | Adjective-group bootstrap 95% lower bound for Efresh > 0 | 0.525535 | Pass |
| 11 | Mean(S+ - S0) > 0 | +0.551279 | Pass |
| 12 | Mean(S- - S0) < 0 | -0.524741 | Pass |
| 13 | Efresh exceeds 95th percentile of absolute matched random-direction effects | 0.538010 > 0.165247 | Pass |
| 14 | Empirical matched random-direction p ≤ 0.05 | 0.009901 | Pass |
| 15 | Qcandidate,fresh exceeds random-direction Q95 | 42.446476 > 17.395144 | Pass |
| 16 | No numerical instability or catastrophic distributional collapse | Pass; max JS 0.038899 < 0.686216 collapse threshold | Pass |
This table is an audit of the frozen confirmatory gate only. Its 14/16 result does not imply a near-pass under the protocol: the rule was conjunctive, so the two failed structural interpretation criteria determine the final confirmatory label.
References
- Tigges, C., Hollinsworth, O. J., Geiger, A., & Nanda, N. (2023). Linear Representations of Sentiment in Large Language Models. arXiv:2310.15154. Primary source.
- Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., & MacDiarmid, M. (2023). Steering Language Models With Activation Engineering. arXiv:2308.10248. Primary source.
- Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., & Turner, A. (2024). Steering Llama 2 via Contrastive Activation Addition. Proceedings of ACL 2024, 15504-15522. doi:10.18653/v1/2024.acl-long.828. Primary source.
- Park, K., Choe, Y. J., & Veitch, V. (2024). The Linear Representation Hypothesis and the Geometry of Large Language Models. Proceedings of ICML 2024, PMLR 235, 39643-39666. Primary source.
- Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O'Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., & van der Wal, O. (2023). Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. Proceedings of ICML 2023, PMLR 202, 2397-2430. Primary source.
- Jorgensen, O., Cope, D., Schoots, N., & Shanahan, M. (2023). Improving Activation Steering in Language Models with Mean-Centring. arXiv:2312.03813. Primary source.
- Lampinen, A. K., Li, Y., Hosseini, E., Bhardwaj, S., & Shanahan, M. (2026). Linear representations in language models can change dramatically over a conversation. arXiv:2601.20834. Primary source.
- Li, S., Wang, Z., Wang, Z., & Li, P. (2026). Uncovering Sentiment Analysis Circuit in Large Language Model. Proceedings of ACL 2026, 20269-20289. doi:10.18653/v1/2026.acl-long.928. Primary source.
- Yao, L. H., Li, Y., & Liu, S. (2026). The Linear Representation Hypothesis Needs a Group Action. arXiv:2609.27158. Primary source.
- Radford, T. J. (2026). Non-Biological Semantics (NBS): A proposed field-level research programme for the study of semantic organisation in artificial systems. Riff Systems Ltd, version 1.0.
Disclaimer & Scope of Application
Non-Biological Semantics (NBS) is a proposed field-level research programme for the study of artificial representational geometry, representational dynamics, functional relevance, and human interpretability. It is not presently an established scientific field or discipline.
This study does not claim that artificial systems are conscious, sentient, subjectively meaningful, or cognitively equivalent to biological systems. The term semantics is used operationally. Structural separation, probe performance, causal steering, or human-readable description must not be treated as interchangeable evidence.
The confirmatory NBS-002 result reported here is STRUCTURAL_PASS_CAUSAL_PASS_INTERPRETATION_FAIL. Exploratory analyses reported after that endpoint do not alter it.
Copyright & Permitted Use
© 2026 Tristan J. Radford. All rights reserved.
Published by Riff Systems Ltd (Company No. 17419314).
This publication may be read, cited, quoted in reasonable extracts, discussed in academic contexts, and used for non-commercial research, provided that appropriate attribution and citation are given and the work is not misrepresented.
Commercial Use
Except where permitted by law, commercial reproduction, republication, adaptation of protected expression, incorporation of protected material into software, products or services, or use of substantial portions of this work for commercial AI or machine-learning training, fine-tuning, dataset construction, or similar commercial exploitation requires prior written permission from Riff Systems Ltd.
Rights Administration
Copyright remains with Tristan J. Radford. Riff Systems Ltd is the publisher and administers commercial licensing and permissions for this publication.
Nothing in this notice limits statutory copyright exceptions or other uses permitted by applicable law. No permission is implied for use of Riff Systems trade marks, logos, confidential material, or unpublished research.
