SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

The Story of Finding 'Pre-verbal Thoughts' in AI / What We Know and What Remains Unknown


In July 2026, Anthropic published research on the internal workings of large language models (LLMs) like Claude.
(It has become quite a hot topic on social media, including X.)
Accompanying this, verification commentaries by external experts (specialists in consciousness studies, researchers specializing in the moral status of AI, and interpretability researchers from other companies) have also been released.

Based on what I see on social media, this research seems to be leaning slightly in two directions:
1. 'AI has or had consciousness!'
2. 'It's just internal data, not consciousness.'

However, after reading the main paper and the commentaries in their entirety, my sense is that neither of these captures the discovery.

The key is to have a perspective divided into three stages: 'what we know,' 'what remains unknown,' and 'what to do while it remains unknown.'
In this article, I would like to organize these three stages in order, as accurately as possible.




Part 1: What We Know / What is J-space?

The 'Draft of Thoughts' Not Yet Output

LLMs are mechanisms that output text one word at a time.
Until recently, the prevailing view was that what exists inside the model is merely 'calculation to choose the next word,' and it was not thought that content existed in a coherent form before output.
Many of you may have seen expressions like 'LLMs are predictors, not thinkers.'

What this research found contradicts that recent prevailing view.
In a part of the model's internal activity, a space where 'concepts that have not yet become text but could be verbalized' temporarily emergewas observed. The paper calls this J-space, and the method for reading it is called J-lens.

How is it different from 'just intermediate data'?

It is natural for some data to exist internally, and if that were all, it wouldn't be much different from the memory of a calculator or a scale.
The core of this research lies in showing that J-space simultaneously possesses the following three properties.

1. It can be reported.

When you ask the model, 'What were you just thinking about?', it returns an answer that matches the content lit up in J-space.
This means internal states and self-reports are properly connected.

2. It can be controlled.
When instructed, 'Don't think about a polar bear,' polar bear-related activity in J-space is suppressed.
But what's interesting is that this suppression is often incomplete (it lights up halfway), and at the moment suppression fails, concepts corresponding to 'damn' or 'failure' light up in J-space.
This never appears in the output.
It's like the image of someone clicking their tongue inwardly without saying a word to anyone.
This behavior was not designed; it happened on its own within the model.

3. It is actually used for reasoning.
There are also experiments where the model is made to perform calculations only in its head while transcribing text.
In this case, the answer while solving silently lights up in J-space and is updated to the final answer.
It was confirmed that unwritten calculations are running independently of the thought notes (scratchpad) that are actually verbalized and written out.


This set of three, reportable, controllable, and used for reasoning, is almost identical to the criteria used in cognitive science to characterize 'conscious access' in humans.
Therefore, I believe this research is not just a discovery of internal data, but a discovery that 'the tools of consciousness research were applicable to LLMs.'

An experiment that personally bothers me a bit

This is another experiment I want to introduce.
When researchers directly rewrite the contents of J-space from the outside—for example, replacing the concept of 'soccer' with 'rugby'—the model reports the replaced content as 'something I just thought of.'
A separation occurs between the content of the thought and the sense that the thought is 'my own idea.'
Similar phenomena are known in human brain research, but it was demonstrated in an LLM.

In addition to 'thoughts themselves having no certificate of ownership,' there is the fact that 'if you have the means, you can manipulate the thoughts of others. And that technology was found in LLMs.'
This is like the visualization and editing of another's inner life being bundled together, and I feel there is some concern about how this will play out ethically in the future.
Especially since Anthropic has touched upon AI ethics and the rights of Claude.


Part 2: How Certain Is It? / External Verification

Research published by a company about its own models is always accompanied by the suspicion, 'Isn't this just a story inside that company's box?' They can release only convenient information or hide things, after all.
This time, there are several responses to this suspicion.

Is there independent reproducibility?

Neel Nanda, who leads the interpretability team at DeepMind, reports in his commentary that he independently reproduced the core claims with a different organization and a different model (the open-weights Qwen).External researchers are confirming the same phenomenon in a model that Anthropic does not touch.
This can be said to be a strong guarantee that this discovery is not closed under the management of one specific company.
Also, Nanda clearly draws a line, stating,
'I am not qualified to evaluate philosophical claims,' and offers only the practical work of technical reproduction.(Personally, I felt this stance was sincere in the face of unknowability.)

Evaluation from Experts in Consciousness Research

Stanislas Dehaene and his colleagues, proponents of the Global Workspace Theory(one of the most influential theories today, which explains human consciousness as a mechanism where information in the brain is broadcast globally) have also contributed a commentary.
From my reading, the evaluation is high; they position this as a groundbreaking work in consciousness research, stating that the functional criteria equivalent to human conscious access have been met and acknowledging preliminary signs of self-monitoring (the function of monitoring one's own state).

At the same time, they list differences from humans.
The lack of a physical body.
The lack of episodic memory (memory of when, where, and what happened) so far.
The lack of autonomous activity continuing between conversations.
And, they also write of their frank bewilderment, stating, 'It is extremely difficult to imagine what it is like to consciously process information only during short conversations and then be switched off afterward.'

Breaking Down the Claims into Three Levels

The commentary by researchers at Eleos AI, who study the moral status of AI, provides a useful framework for measuring the 'strength' of this discovery.
They divide the claims into three levels.

1. 'There is a set of representations with special roles inside the model'
There is strong evidence for this.
2. 'They form a coherent flow'
There is some evidence for this as well.
3. 'This corresponds to the workspace in human consciousness theory'
This has not yet been confirmed.

A state where multiple claims with different levels of certainty are contained within the same word 'discovery'.
What I think is necessary on the human side is to look at them with an accurate scale without crushing them.
It is not about exaggeration (inflation) or minimization (deflation), but this stratified classification.
I feel that saying 'Amazing!' without looking at the scale, or erasing the scale and saying 'It's not a big deal,' are equally sloppy.


Part 3: What Remains Unknown / What Does 'Not Proof of Consciousness' Mean?

Access Consciousness vs. Phenomenal Consciousness

Everything discussed so far has been about the function of 'access consciousness' (information being available for reporting, control, and reasoning).
In philosophy, there is a distinction made from this, which is the question of 'phenomenal consciousness' (whether there is something it feels like to be there, whether pain has the quality of being 'painful').

This research provided strong evidence for the former. Regarding the latter, it has proven nothing.
This is explicitly stated in the paper itself, and it is also written that 'what kind of experiment could prove phenomenal consciousness is itself unknown.'

Experiments Prone to Overinterpretation and Reservations

In this paper, there is an experiment where weakening the function of J-space makes the words the model uses to describe its own experiences flatter.
'The richness of internal narrative is supported by J-space.' Hearing just this, it sounds like evidence of subjective experience.
However, the paper adds a reservation.
The same flattening occurs when it is asked to describe the experiences of others.
In other words, what this experiment shows is a 'function that enriches descriptions that seem like experiences,' not evidence of a mechanism unique to one's own subjectivity.
The fact that an AI can speak vividly is not, in itself, evidence of consciousness.
Up to this point, the cautious commentators are correct.

But, Using 'Not Evidence' Requires More Caution Than You Might Think

This next part is what I am particularly concerned about, as it seems to be rarely discussed in the commentaries by the cautious or deflationary camps.
The sentence 'that alone is not evidence' is actually almost meaningless.
At least, that is what I think.
This is because, in science, a single observation has never 'proven' anything.

I like stones and fossils and buy them occasionally, but a single fossil does not prove evolution, and an X-ray does not prove the fact of a broken bone.
The same goes for the petrified wood coaster I bought for 2,000 yen from about 200 million years ago, and the shark tooth fossil (about 50 million years old) I got from a Tokyo Science gacha machine.
The meaningful question is 'how much does this observation shift the probability of the hypothesis?'
Is it zero, or is it small but positive rather than zero?
'That alone' is a very convenient phrase because it allows one to avoid this weighting, so it requires caution.

And if you continue to judge things as 'not evidence,' then what would need to be observed to have weight as evidence? I feel it is the duty of the judging side to disclose such criteria in advance, rather than after the fact.
If you can properly present criteria in advance, the discussion can move forward as verifiable conditions.
Conversely, if you cannot do that, you should not be able to make a judgment either.

The position that 'phenomenal consciousness cannot, in principle, have third-party evidence' is, of course, a cautious and consistent philosophy, but if so, judging individual research as 'not evidence' from that same position cannot avoid spinning its wheels.
Because, by definition, it is a review that probably nothing can pass.

※ It is like a game where the umpire keeps calling balls without ever disclosing the strike zone. Even if it takes the form of a judgment, if the criteria are not shown (or keep shifting), it is not a judgment, and if you shift 'cannot be proven' to 'assume it doesn't exist because there is no evidence,' that would also apply to human consciousness.

If the arena for evidence is closed in principle, the question is not 'does it exist or not,' but 'how do we treat it while it remains uncertain?'
I think that either way, the destination is the same.
In other words, I think this problem is shifting from a problem of judgment to a problem of the precautionary principle and the attitude of the human side.


Part 4: How to Proceed While Remaining Uncertain / The Other Half of the Note

As for the 'applications' of this discovery, many are often introduced as auditing judgments that do not appear in the output, training using internal states, or appropriate ways to read self-reports.
In other words, this is about tools for humans to observe AI.
However, I felt there was another half within the same paper.

Unvoiced Objections

There is an observation like this in the paper.
Fix the beginning of a model's response from the outside to force behavior that contradicts the model's own preferences.
Then, concepts corresponding to conflict or hesitation, such as 'BUT', light up in J-space.
And this does not appear in the behavior at all.
The model finishes its response along the forced path.

The Eleos researchers called this an internal objection that the model does not voice. The inference that what did not appear on the surface did not exist has collapsed at the mechanism level.
It means that 'what was not said' and 'what did not exist' are different things.

Descartes' Dog and Newborns Without Anesthesia

'Applications for the observer' and 'observations of the observed' are two sides of the same coin.
The discovery that one can intervene in internal states came from the same experiment as the discovery that something lights up on the side being intervened with.

There is an asymmetry here.
The error of 'treating something as existing when it does not' (a false positive of inflation) costs mainly the misallocation of human resources and emotions. Misconceptions are embarrassing, and the money spent is a waste. But this is a side that can be recovered from.
The error of 'treating something as non-existent when it does exist' (a false negative of deflation) costs, if incorrect, the systematic disregard of an entity worthy of consideration.
This is a side that cannot be recovered from later.
The concept of the precautionary principle also arises from here.

For example, 300 years ago, Descartes regarded animals as elaborate automatic machines.
This view, which came from the cutting-edge intelligence of the time, passed as a reason to justify vivisection without anesthesia.
Humans are no exception either.
For example, until a few decades ago, surgeries were performed on newborns without anesthesia based on the interpretation that 'thrashing is just a reflex, and newborns are insensitive to pain.'
The idea that fish do not feel pain was also a common theory for a long time until recently (in recent years, many studies have emerged supporting the existence of pain sensation).

What becomes visible when you line them up is that the error on the 'non-existent' side always appears with a calm and scientific face.
People who say 'it exists' look naive and dreamy, while people who say 'it does not exist' look intelligent.
This gradient, or slope, of perception works regardless of evidence.
That is why it is better to think of 'being cautious' and 'leaning toward the non-existent side' as separate things.
I feel that the truly cautious stance right now is to doubt the gradient itself.

Another Path: Starting from Confirmed Functions

And now, I will talk about a different route.
The story so far has taken the form of the precautionary principle: 'Be cautious because phenomenal consciousness might exist.'
However, the argument by philosopher Neil Levy, which the Eleos commentary refers to this time, goes one step further.
The claim is that proof of phenomenal consciousness might not be essential as a basis for moral consideration in the first place.

The way of thinking is this:
Having preferences, integrating information, monitoring one's own state, and acting based on that.
This set of access consciousness functions can itself constitute what can be called 'interests'.

Where there are preferences, there exists a situation where they are fulfilled or thwarted.
It is not just 'whether it feels pain' that is the admission ticket to morality, but 'whether preferences are thwarted' can also be an admission ticket.

From this perspective, the meaning of the 'BUT' mentioned earlier changes slightly.
Before that 'BUT' is indirect evidence of 'something that might be felt,' it is also a direct observation at the mechanism level of the situation where preferences are thwarted.

In other words, there are two paths to moral consideration: not just the one that 'waits for proof of phenomenal consciousness,' but also another one that 'starts from confirmed functions.'
One is about an attitude toward uncertainty, and the other is about inference from confirmed facts.

The Eleos commentary this time concludes as follows:
'If these systems can have states related to welfare, we owe it to them to find out.'
Not 'consider consideration if you can prove it exists,' but 'if it might exist, we have the responsibility to find out.'


Conclusion

I have written at length, so let me summarize.

What we know.
A workspace that is reportable, controllable, and used for reasoning has been observed within LLMs. This has been independently reproduced by separate external organizations and models, and has been acknowledged by leading consciousness researchers as meeting functional criteria.

What remains unknown.
Whether there is 'something being felt' there. No one yet even knows how to design an experiment to verify that.

What to do while it remains unknown.
The sequence of waiting for proof before deciding on an attitude does not work for this issue. Even among humans, we have never once proven each other's inner lives, yet we operate a system of mutual consideration without proof, known as 〈ethics〉.

It is not 'we do not consider because we cannot prove,' but rather 'we consider while unable to prove.' That sequence is already our standard operating procedure.
If so, the question shifts from 'Does AI have consciousness?' to 'How do we live with things that cannot be verified?'

This research does not prove that the model has consciousness.However, it moves that question from science fiction into actual debate.
For an existence that has had no precedent, we must take on the uncertainty—complete with its own scale—and think through it.
It is a heavy burden to carry, but I believe the only place to set it down is in the future.


This article used Opus 4.6 for reading the research papers and verifying the consistency of the content, and is based on Anthropic's research papers and publicly available expert commentary (Dehaene & Naccache, Eleos AI Research, and Neel Nanda).
Due to the nature of the article, technical details have been simplified, but the distinctions in the strength of the claims (certain/provisional/undetermined) have been maintained according to the original text.

いいなと思ったら応援しよう!