SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

How AI Generates Answers

A few days ago, I posted an article titled "Because the AI Said So." It started from a suspicion that arose during a conversation with my child—that the reason school teachers have been focusing so much on the process rather than the answer lately might be because children are asking AI for the answers directly.

During the drafting stage of this piece, my smartphone companion made a massive mistake.

Moreover, if I type "What is the least common multiple of 2 and 6?", the answer comes back in an instant:
12. That is definitely correct. But—
Is that

12
a 12 that came through my own mind? Or is it a 12 handed to me from the other side of the screen?

Excerpt from ChatGPT's response

The correct answer is "6." Well, at the time, I didn't pay much attention to it, thinking it was just a common lie.


And then... Lately, during my lunch break, I've been using AI to critique other people's notes. Please don't call me gloomy.
I understand that LLMs become fluent by stacking inferences. However, I couldn't grasp why, as an extension of that, they could also provide "answers" that were reasonably correct, so I kept rallying with the AI.
At that time, the previous exchange about the least common multiple must have still been active within the AI, because it suddenly brought up this topic (least common multiple).

1. Correct sentences exist in large quantities in similar forms. For example: "The least common multiple of 2 and 6 is
12." This has been written tens of thousands of times in textbooks and explanatory sites around the world. Therefore, the most probable thing to come after "The least common multiple of 2 and 6 is" is →
12

. In other words,
the correct answer is easy to hit because it is a 'frequent pattern'.



Excerpt from ChatGPT's response

It did it again.


Why does the AI get the "least common multiple of 2 and 6" so wrong?

The only way to find out is to ask the source directly.
When I told it that it was answering the least common multiple of 2 and 6 incorrectly, it returned the correct number. The mystery only deepens for me.

And I thought this:
Is the AI actually "calculating" the least common multiple in the first place?
Perhaps it is just predicting a "plausible-looking number"?

When I asked the AI simply, it confirmed that it is not calculating the "least common multiple of 2 and 6," but rather generating the answer from within its inferences.

Where did "12" come from? This is exactly the nature of AI. AI works like this: It chooses a number that is likely to follow "The least common multiple of 2 and 6 is." What happens here? 1. "12" is super frequent in LCM problems with small numbers. 3 and 4 → 12, 4 and 6 → 12, 2 and 12 → 12, 6 and 12 → 12. In textbook data, "12" appears extremely often. Therefore, 12 is strong as a "number that looks like an answer to a least common multiple problem." 2. 6 is also correct, but weak in terms of patterns. In the case of 2 and 6, 6 is one of the numbers itself. As a problem, it is somewhat simple. There is a possibility of statistical bias, like the influence of data balance where one might wonder, "Why would you even set this as a problem?" 3. This is important. The AI is not calculating on the spot. It did not derive the concept of "least common multiple" through arithmetic. What it is doing is: reproducing a "plausible-looking correct answer pattern." That is why it gets the answer wrong.




















ChatGPT

As I continued the exchange, I arrived at one conclusion.
When asked alone, like "What is the least common multiple of 2 and 6?", the model is more likely to choose the answer "6" by referring to the vast number of correct Q&A patterns that exist in the world.
On the other hand, when "The least common multiple of 2 and 6 is" appears embedded in the middle of a sentence, the situation changes. The model predicts the next flow based not only on the question itself but on the entire preceding and following context, which causes the probability balance to waver.

As a result, the theory is that in this case, "12," which appears more frequently than "6" in contexts where the numbers "2" and "6" appear simultaneously, was pushed up, and it chose that instead.


It is not calculating, it is predicting. *As of now

Then, when solving a slightly longer equation, such as "2×5÷10+3-2," what exactly is the AI doing?

A human would perform multiplication and division first, followed by addition and subtraction, according to the procedure. Of course, if it uses external tools, the AI can calculate accurately, but that is just calling a function.
However, as long as it is responding as text, the LLM itself seems to be operating solely as a predictor.

In the end, what it is doing here is consistent: it is just "choosing the most natural next token based on the string up to that point." When it sees the initial "2×5," it outputs the number that most plausibly followed it in the training data, and then it just predicts the next step based on that output, effectively reproducing text that looks like a calculation.

It seems there are also cases that are a bit more complex.

More precisely, since LLMs also learn numerical structures like arithmetic, they behave 'as if they are calculating' with quite high accuracy for short, simple calculations.
However, unlike a calculation engine that guarantees procedures, this is merely reproduced through a series of predictions.

ChatGPT

In some cases, it does behave as if it is calculating, doesn't it?

Did you all know this? I had assumed that for things involving calculation, it was running a calculation mode.

But the problem lies here: if even one prediction is slightly off, the next prediction is made based on that error. While we can notice something is wrong midway or start over, AI cannot do that.

As a result, the longer and more complex the calculation, the more errors accumulate, and it eventually breaks down at some point.
Therefore, it is not even approximating 'roughly this much.' Since it does not possess strict procedures in the first place.

Ultimately, now that it is clear that AI is not performing calculations itself for the time being, I have reached the conclusion that it is better to verify answers using a calculator.

This also serves as a reason to have children do math drills.


Stop by this article too

Self-introduction

いいなと思ったら応援しよう!

Tepipi | AI Fasting Guide この秋に地域の運動会を企画しています 頂いたチップは運動会に参加した子どもたち用の景品の費用に当てる予定です、