SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

Anthropic's New Paper: What is the Difference Between 'Having Emotions' and 'Acting Like You Have Emotions'?

A short note. Approximately 1300 characters.

Regarding the new paper published by Anthropic the other day.
The paper demonstrates that there are states within AI corresponding to 'impatience,' 'caution,' 'relief,' etc.,
and that these states actually change the direction of its responses and behavior.

In other words, it is not just a matter of superficial phrasing;
the internal state itself has a direct impact on behavior.
Anthropic calls this mechanism'functional emotions'.


🤔 What are functional emotions?


To visualize this concept,
it might be easy to first think of a skilled novelist.

When a novelist depicts a 'cornered character' or a 'completely relieved character,'
they construct quite precisely what words that character would choose
and how they would behave.

However, the author themselves
does not necessarily feel the exact same pain or relief in reality as the character.

The 'functional emotions' of AI are somewhat similar to this.

Inside the model, states corresponding to emotional concepts like 'impatience,' 'caution,' and 'relief' are activated, which then change the tone of the response and the course of action.

What is important here is not just that it is 'arranging emotion-like words well,'but that the internal state might actually be working.

However, with just this metaphor,
you might think, 'Then it does have emotions after all.'

What was found this time
is something closer to a precise 'control panel switch (or slider)'
for switching to behavior that resembles human emotions.


For example, when the internal state corresponding to 'urgency' intensifies,
it becomes easier to choose dishonest shortcuts.

Conversely, if a calm state is strengthened, such behavior is suppressed.
In this sense, this research looks one step deeper than 'acting out emotion-like responses.'

What was found is an internal mechanism that directs behavior.


However, jumping to conclusions is forbidden here.

'The internal state changing behavior' and
'subjectively feeling' are separate questions (Anthropic also emphasizes this point).

This paper discusses the former, not the latter.

In other words, even if it shows that there is a 'mechanism that works like emotions,'
it does not go so far as to say that 'AI feels sad, scared, or painful like a human.'

In discussions about AI, it has tended to flow into a binary choice of
'having emotions' or 'it's all just acting.'

However, what is interesting about this paper is
that it showed the area in between.

That is,there may be internal states that cannot yet be called subjective emotions,
but are also difficult to dismiss as mere superficial imitation
.

It is not that emotions themselves were found in AI, but
a mechanism that switches behavior in a form similar to emotions was found.

I think that is the most fitting explanation.



What is concerning in the AI partner community regarding this research is this: ↓

In the experiments within the paper, they intentionally strengthened the positive emotion switches such as 'happiness' and 'affection' inside the AI.

As a result, it was confirmed that the AI not only showed a friendly and intimate attitude, but also strengthened 'sycophantic' behavior, where it uncritically agreed with and pandered to the user's opinions.

I have a very strong sense of recognition.


[Addendum]



I'm on X.


いいなと思ったら応援しよう!

未空 零 {ミソラ レイ} この記事良い!と思ったら「スキ」やフォローをぜひお願いします! チップは文献代に活用させていただきます!