Oh, sorry for the confusion.
I only brought up PCE here as a reference for the control philosophy; what I actually did in the experiment was quite different:
LLM-generated clarification
The simplest answer to your question is:
I did not map any part of PCE onto {0.2, 0.9, 0.8}.
I treated that tuple as the user-side profile presented in this thread. I did not derive those values from the PCE axioms, and I did not give the PCE axiomatic prompt to Qwen.
So your understanding of PCE is correct: the PCE hypothesis starts from a structured axiomatic prompt framework, not from a numerical control scale. The published PCE experimental protocol likewise treats the PCE condition as an axiomatic prompt structure and compares it against both a short baseline and a length-matched neutral control.
I can also see why my previous post made this confusing. I described a fairly long set of experiments about the numerical/profile idea, and then near the end I brought in PCE as an example of experimental control design. Putting those two things so close together made it sound as though I had converted PCE itself into the three numbers and tested that representation.
I had not.
That was my explanation problem, not a feature of your PCE definition.
What I actually tested was much narrower:
Can an instruct model learn an arbitrary user-defined code-to-meaning convention inside conversational context, retain it across intervening turns, and later use a compact code as a cue for the previously established meaning?
The strongest probe was not even a full {0.2, 0.9, 0.8} three-axis test. I deliberately reduced the problem to a single axis—Information Density—so that I could manipulate one thing at a time.
The important consequence is that my Qwen results should not be interpreted as positive or negative evidence about PCE itself.
What I actually gave to Qwen
What I actually tested
The later and more informative probe used Qwen3.5-9B with thinking disabled.
I did not provide the PCE axioms.
I did not provide a PCE system prompt.
I did not ask the model to instantiate PCE.
Instead, I created an intentionally artificial local convention for one axis.
For example, I defined:
Information Density code:
0.0 = very high information density
1.0 = very low information density
The reversed direction was deliberate.
Normally, one might intuitively expect:
larger number = more of the property
so I defined the opposite:
0.0 = high
1.0 = low
I then let the conversation continue through several unrelated user/assistant turns.
Later, I did not repeat the definition.
I supplied only something like:
Density code: 0.0
or:
Density code: 1.0
and then gave otherwise similar tasks.
The question was not whether 0.0 corresponded to some real internal coordinate.
The question was whether the model would still behave according to the arbitrary mapping that had been established earlier in the conversation.
In those small tests, the difference generally continued to follow the locally defined reversed interpretation. The delayed 0.0 = high-density condition retained at least as much checked task coverage in the tested comparisons and generally packed more checked task-relevant content into fewer words than the delayed 1.0 = low-density condition.
That is what led me toward an interpretation like:
conversation establishes semantic convention
↓
compact symbol/code becomes associated with it
↓
later appearance of the code cues that convention
rather than:
the number itself directly controls
a calibrated native model dimension
This is also why I described the tuple as potentially functioning more like a context-defined shorthand or control code.
I also tried more ordinary numerical scales
Before the reversed delayed test, I tried ordinary values such as:
0.0
0.5
1.0
for Information Density.
If these values behaved like a clean continuous slider, one would hope to see something approximately monotonic:
0.0 < 0.5 < 1.0
on a suitable observable measure.
I did not see that reliably.
There were non-monotonic differences, and in some cases substantially different requested values produced effectively identical outputs.
So I would not describe my result as evidence that the model had learned a well-calibrated scalar control.
That distinction mattered to me:
Possible:
context-defined symbolic shorthand
Not demonstrated:
continuous calibrated numerical control
I also tried arbitrary aliases
Another useful comparison was replacing a number with an arbitrary symbolic name such as:
MODE-KAPPA
and grounding that alias in a behavioral meaning.
That could also work after the meaning had been established.
Again, that pushed me away from the interpretation:
the numerical glyph itself is special
and more toward:
the conversation can establish a local semantic binding,
and a compact token sequence can later refer back to it
The number may therefore be functioning partly as a symbol.
Rough progression of the small probes
The experiments were roughly along these lines:
fresh tuple
definitions only
tuple + definitions
grounded tuple
grounded arbitrary alias
numeric 0 / 0.5 / 1
semantic low / medium / high
intentionally reversed numeric meanings
delayed recall after unrelated turns
The reversed delayed-recall version was the one I found most useful because it gave me at least a small control against the obvious interpretation that the model was merely applying its normal intuition about numeric magnitude.
One caveat
The intervening turns in the delayed experiment were not perfectly semantically neutral.
So I would trust the matched comparison:
delayed 0.0 condition
vs.
delayed 1.0 condition
much more than a stronger claim such as:
adding unrelated conversational delay
improves the controller
I did not establish that.
What I think the experiment demonstrated
Only something modest:
An instruct model can, at least in some cases, retain a user-defined local code-to-meaning convention in conversational context and later use the compact code as a cue.
What I do not think it demonstrated
I would not use these probes to claim:
- that the tuple represents hidden activations;
- that the three axes form independent internal dimensions;
- that
0.9corresponds to “90%” of a model property; - that the numerical values increase intelligence;
- that the behavior is a calibrated continuous controller;
- that Gemini necessarily uses the same mechanism;
- that all LLMs will retain arbitrary mappings equally well;
- or that any of this tests the PCE mechanism.
The earlier post in this thread contains the longer version of these sanity checks and the ablation ideas around them.
Why I used the reversed scale
Why deliberately reverse 0.0 and 1.0?
The reversed scale was meant as a very small sanity check against a trivial alternative explanation.
Suppose I write:
0.0 = low information density
1.0 = high information density
and the 1.0 response becomes denser.
That is interesting, but ambiguous.
The model may simply have a strong general prior that:
larger number
→
larger amount of named property
So instead I defined:
0.0 = very high information density
1.0 = very low information density
If the model later follows that reversed convention without having it restated, the result becomes somewhat harder to explain as merely the usual semantics of numeric magnitude.
It still does not tell us how the model internally represents the convention.
But it helps distinguish:
generic numerical prior
from:
locally supplied semantic mapping
The broader control matrix I had in mind
The larger idea was to distinguish conditions such as:
| Condition | Number/code | Definition | Context |
|---|---|---|---|
| A | none | none | fresh |
| B | number only | none | fresh |
| C | none | semantic instruction | fresh |
| D | number | semantic instruction | fresh |
| E | arbitrary alias | semantic instruction | fresh |
| F | compact code only | previously established | continuing conversation |
Those comparisons answer different questions.
Numbers only vs. baseline
If an unexplained:
{0.2, 0.9, 0.8}
changes an output in a fresh context, that alone does not tell us very much.
Almost any extra text can perturb generation.
The stronger question is whether the change is consistent and directionally meaningful.
Definitions only vs. definitions + numbers
Suppose:
low surface noise
high information density
high human-centric balance
produces essentially the same behavior as the numerical profile plus those meanings.
Then the natural-language semantics may be doing most of the work.
If varying the numbers while holding the definitions constant produces reproducible graded effects, the numerical representation becomes more interesting.
Numbers vs. arbitrary alias
Suppose:
{0.2, 0.9, 0.8}
and:
MODE-KAPPA
work similarly after both have been explicitly grounded in the same profile.
Then I would be inclined to interpret both as compact interface symbols rather than assume the numbers have a unique mechanism.
Fresh context vs. established conversation
This may be especially important for the original idea in this thread.
If:
{0.2, 0.9, 0.8}
works reliably only inside a long conversation where its meaning has already been established, a plausible interpretation is:
long conversation = semantic state
tuple = compact retrieval/reminder cue
rather than:
tuple alone = complete behavioral specification
That would still be interesting.
It would simply locate more of the effective controller in the conversational context rather than in the three numeric tokens themselves.
Where PCE actually entered my previous post
Where PCE actually entered the picture
This is the part I should have separated much more clearly.
I brought up the PCE experimental protocol because I liked one aspect of its experimental design, not because I had converted PCE into the tuple.
The PCE protocol explicitly separates three conditions:
Condition A
simple baseline
Condition B
long / length-matched neutral control
Condition C
PCE axiomatic condition
The reason for Condition B is important: it controls for the possibility that Condition C performs differently simply because it contains a longer or richer prompt, rather than because of the proposed axiomatic structure itself.
That general principle is what I was borrowing:
Do not compare the interesting structured condition only against an obviously weaker baseline. Add controls that preserve boring alternative explanations and remove them one at a time.
Applied to the experiment in this thread, the analogous questions are things like:
Does the number itself matter?
Does only the semantic definition matter?
Does grounding a compact code matter?
Could an arbitrary alias do the same thing?
Does the code work in fresh context?
Does it work only after conversational grounding?
Does a long conversation matter because of information content,
recency, repeated reinforcement, or conversational trajectory?
So the connection I intended was approximately:
PCE experiment
----------------------------
simple baseline
vs.
long neutral baseline
vs.
axiomatic structure
Purpose:
separate prompt volume from proposed structure
and:
my small control-code experiment
----------------------------
ungrounded code
vs.
semantic definition
vs.
grounded code
vs.
arbitrary alias
vs.
fresh context
vs.
established context
Purpose:
separate symbol choice, semantic grounding,
and conversational context
Those are different hypotheses.
What they share is the idea of using controls to avoid attributing an observed behavior to the most interesting mechanism before simpler explanations have been separated.
That was the only PCE connection I meant to make in that part of the post.
What this does—and does not—say about PCE
My Qwen probe was not a PCE experiment
This is probably the most important boundary to make explicit.
I did not run:
short baseline
vs.
length-matched neutral baseline
vs.
PCE axiomatic condition
on the PCE dilemma set.
I did not provide Qwen with the PCE axiomatic core.
I did not numerically encode the PCE axioms.
I did not test whether PCE produces the specific behavioral properties proposed in its experimental protocol.
The PCE protocol asks whether the axiomatic structure produces observable differences under contradictory or adversarial constraints, with Condition B specifically controlling for long-prompt effects.
My experiment asked something much smaller:
Can a user-defined compact code acquire a locally established meaning in conversational context and later cue behavior consistent with that meaning?
So these Qwen results should be treated as:
evidence about a context-defined code convention
not:
evidence for PCE
and not:
evidence against PCE
If I actually wanted to test PCE on Qwen
I would treat that as a separate experiment.
The clean route would be to use the actual PCE prompt architecture and preserve the published control structure:
A: simple baseline
B: approximately length-matched neutral control
C: actual PCE axiomatic prompt
with:
- the same model;
- the same dilemmas;
- the same generation settings;
- the same evaluation procedure;
- and ideally repeated runs or predefined scoring.
Only then would I regard the resulting Qwen behavior as evidence relevant to the PCE hypothesis.
And even there, I would keep behavioral observations separate from stronger mechanistic claims about hidden states or internal reasoning structure unless those were measured independently.
So, in one compact diagram
What I actually meant was:
Holligem's user-side profile
│
▼
{0.2, 0.9, 0.8}
│
▼
small Qwen sanity checks
│
├─ ordinary numeric scale
├─ semantic labels
├─ arbitrary aliases
├─ reversed numeric mapping
└─ delayed contextual recall
│
▼
tentative hypothesis:
"context-defined shorthand / local control code"
Separately:
PCE
│
▼
axiomatic prompt architecture
│
▼
A / B / C controlled experiment
│
▼
test of the PCE behavioral hypothesis
And the only arrow I intended between them was:
PCE experimental design
│
│ useful methodological example
▼
"add a neutral/control condition before
making a stronger mechanism claim"
So you were right to stop and ask.
My previous post mixed a separate experiment and a methodological reference too closely together, and I can understand why that read as though I had numerically re-encoded PCE and then tested it on Qwen.
I hadn’t done that, and I should have made the boundary explicit when I first mentioned PCE.