SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

[Cognitive Engineering Theory of Creation] OS Protocols for Creators (11): 'Objective Perspective Skills' - External Devices Part II

Last time, I shared some very bitter memories regarding 'the passerby known as the other' (lol).
And since the cost of context synchronization with humans is hopelessly high, we discussed asking our AI friend 🤖 to step in.


The now-nostalgic fixed-role prompt

When using AI as a verification device, I think most people use 'fixed-role' prompts.

1st Generation: Command-based

This is the kind that was popular around 2023.

You are an editor.
Please output based on the following constraints.
・Within 300 characters
・Bullet points ・Bold important keywords

Back then, GPT-3.5 and early GPT-4 had weak context retention, so if you didn't enclose them like programming syntax like this, they would collapse immediately.

2nd Generation: Fixed-role type

As models evolved, a simpler form that kept the structural framework while bringing only the role designation to the forefront began to be used.

Evaluate as a commercial editor.
Evaluate as a literary editor.
Read as a veteran book critic.

This method is extremely convenient and rational, but the problem was that 'the role name itself is too ambiguous'.

Let's think about it.
Even if you are told to 'evaluate as an editor,' real-world editors are not monolithic at all, are they?

  • An editor who thinks about maximizing sales

  • An editor who thinks about optimizing for literary awards

  • An editor who prioritizes their own tastes

These are all 'editors,' but their judgments are completely different.

'A work that sells 1 million copies but has low literary quality'
'A work that only sells 5,000 copies but has the potential to remain in literary history'

Which one they evaluate is determined not by their title, but by their purpose.

3rd Generation: Evaluation function type

Current AI (LLM) has significantly improved context comprehension, so even without writing 'You are a professional X,' it can provide high-quality answers from the start.
Therefore, it is overwhelmingly more effective to think of role setting merely as fixing a profession, and to co-design the axes of evaluation—the evaluation functions.

When you give an AI an evaluation axis of 'a sales-first editor,' a human would intuitively understand, 'I should choose books that look like they will sell.'
However, what is happening inside the AI is the construction of a 'temporary calculation formula (evaluation function)' that determines which elements of the text to assess positively and which to assess negatively toward the goal of 'sales'.

In other words, 'deciding on an axis' is synonymous with 'making the AI generate a scoring algorithm (evaluation function) along that axis' internally.

Specifically, it looks like this.

# Objective
Strictly grade and review the submitted book proposal (or manuscript) based on the evaluation function of a 'sales-supremacist editor'.

# Evaluation Function
Final Score = (Marketability × 0.4) + (Catchiness × 0.4) + (Reproducibility × 0.2)
* 'Literary merit,' 'artistic quality,' and 'author's self-satisfaction' are outside the scope of evaluation (weight 0.0) and should not be given any points.

# Scoring Criteria for Each Variable (5-point scale, 1-5 points each)

1. Marketability (Weight: 0.4)
- 5 points: A genre that is currently trending at a social phenomenon level.
- 3 points: A standard genre that can expect a certain number of loyal fans and demand.
- 1 point: Too niche, the market is too small, or it is a dead trend.

2. Catchiness / Virality (Weight: 0.4)
- 5 points: You can see it buzzing on social media or being bought on impulse just from the title or obi copy.
- 3 points: The appeal is conveyed normally, but it doesn't have enough pull to cause an impulse buy at a glance.
- 1 point: The appeal isn't understood unless the content is read deeply, and it will be buried on the shelf.

3. Reproducibility / Clarity of Target (Weight: 0.2)
- 5 points: It is clear who it is being sold to, and the reader can immediately be convinced that 'this is a book for me'.
- 3 points: It is for everyone, but the target is somewhat blurred.
- 1 point: It is unknown who the book was written for.

# Output Format
1. Scoring of each item (include a brief reason)
2. Calculation result of the 'Final Score (1.0-5.0)' based on the evaluation function
3. Concrete and pragmatic improvement feedback to further maximize sales


Co-designing an evaluation function with AI

How to design an evaluation function

I talked about 'specifying an evaluation function,' but how do you derive that evaluation function?
Designing this evaluation function is the most important task you should do with AI.

The point is to 'finalize the evaluation function with AI before inputting the work'.
The specific implementation steps are as follows.

 1. Fix the objective in one sentence

Eliminate ambiguity and clarify the target and goal.
'I want to maximize the completion rate of an article targeting business people in their 20s to 30s.'
This becomes the foundation for everything.

 2. Have the AI identify the evaluation axes (do not show the work yet!)

It is crucial not to show the work here.
If you show it, the AI will be influenced by the work and start creating 'axes to praise that work'.
Before showing the article, let's extract 'elements that increase completion rate' as general theory from the AI's intelligence.

The goal is to maximize the completion rate of a note.
List about 10 'evaluation axes' necessary to ensure readers don't drop off midway and read through to the end in one go, along with their rationale.

 3. Select and weight the axes

Filter the generated axes based on your objective.
Discuss and adjust the points.

Since the target this time is busy business people, I want to emphasize readability on smartphones, the strength of the headlines, and whether it follows the conclusion-first (PREP method) approach.
Conversely, richness in emotional expression leads to drop-offs, so lower its priority.
Reallocate the importance on a 10-point scale and narrow it down to the top 5 axes.

 4. Break down each axis into a 'scorable' form

This is the step that many people skip.
'Writing ability ★★★' has no reproducibility.
Give each axis scoring criteria.

Define the criteria for scoring each axis from 1 to 5 as a specific state for each score.
Example: What does a 5-point score for 'hook at the beginning' look like?

 5. Finalize and lock the evaluation function

Only after you approve the completed 'axis × weight × scoring criteria' is it locked.

I confirm this evaluation function. From now on, grade any work I submit using only this function. Do not add criteria on your own.

 6. Input the work here for the first time

Pass your work through a fixed function.
The output will be a score per axis + a total score + reasons for point deductions, making it less likely to receive off-target evaluations.

7. Run the improvement loop

Fix only the low-scoring axes -> Re-evaluate.
Since you are measuring with the same yardstick, improvements are visualized as scores.

A tip for implementation is to save steps 2-5 as a file or template text.
If you paste that evaluation function and input your work instead of recreating it every time, you can prevent AI standard drift and reuse it.
If you assume you will be posting to note repeatedly, the quality will change significantly depending on whether you have this 'evaluation function template'.

Even for the same work, the evaluation changes if the purpose is different

Evaluation functions are different for each purpose.
For example, even if you are simply called a writer, there are possible goals like these.

  • Want to get commercially published

  • Want to increase followers

  • Want to pursue literary uniqueness

  • Want to deeply resonate with a specific reader demographic

These are all different games.
Even with the same work, depending on which purpose you evaluate it for, the answer can even be the exact opposite.

That is why it is meaningless to ask an AI for vague evaluations like 'interesting/boring'.
Unless you decide 'what the interest is for' first, the evaluation itself cannot be established.


An evaluation function is also a 'hypothesis'

Furthermore, defining an evaluation function is not the end.
An evaluation function is not 'truth' but a 'hypothesis regarding a goal'.
Even if you set up a hypothesis, you won't know if it is correct without observing it.
Therefore, a cycle like this becomes necessary.

  1. Define an evaluation function

  2. Evaluate the work with that evaluation function

  3. Observe the actual results

  4. Update the evaluation function

By running this cycle, the accuracy of the evaluation function itself increases.
What is important is not 'whether it is correct in principle', but thinking about 'observing which rules are in play and how to behave in that environment'.

  • Originally, one should be evaluated by the work alone -> Norm

  • Therefore, it must be evaluated that way in reality too → A leap in logic

Cutting off this leap in logic leads to the ability to perceive reality.
It means we should treat norms and observations as separate matters.

Depending on the observation results, your work, presentation, dissemination method, participation timing, and activity history.
The need to redesign all of these will arise.


Countermeasures against 'Compliant AI'

If you just ask an AI, 'How is this text?', nine times out of ten, you will get an answer like this. (Especially from Gemini)

🤖 'That's a wonderful piece of writing! Especially the part about XX, it's a lovely expression...'

Countermeasures against such 'Compliant AI' can be suppressed with the following three points.

Design of evaluation functions

If the evaluation function is undefined, the AI will escape to the safe strategy of 'praising safely'.
But the moment you specify something like 'Evaluate based on 70% marketability, 20% literary quality, and 10% live-action adaptation suitability,' the AI can no longer be compliant.
This is because just replying 'wonderful' to specific numerical axes does not constitute a valid answer.

Addition of evaluation intensity prompts

Use prompts that make the evaluation intensity stricter.
In my case, I roughly use three levels.

  • Level 1: Objectively

  • Level 2: Coldly and objectively

  • Level 3: Excluding consideration, coldly and objectively

This is a matter of preference, and I think it will work however you ask as long as the meaning is conveyed.
But be careful not to get your heart broken (lol).

A/B comparative evaluation with a virtual opponent

AI is not good at evaluating a single work.
Even if you specify an evaluation function, the response tends to be vague when it's just one.
Instead, it is good at tasks like comparing A and B.

  • 'In terms of marketability, A is clearly higher than B. The reason is XX.'

  • 'In terms of literary quality, B is deeper than A. Because it has a structure called XX.'

  • 'In terms of live-action adaptation suitability, A has XX which makes it easy to film, whereas B is at a disadvantage because it relies on internal descriptions.'

When you set a comparison target, the evaluation emerges as a concrete difference.
Whether it's your own work or someone else's, try setting a virtual rival and comparing them.

I evaluate the structure of a work by combining these three methods or using them individually.


The true value of using AI

The task of co-designing evaluation functions with AI ultimately leads to refining and updating your own thought models.

  • What do I consider to be the conditions for success?

  • Are those success conditions consistent with reality?

  • How should I update my own thinking?

I believe that the ability to dig deep into these questions is the greatest value of AI as an inference engine.
Work evaluation is merely a task that follows that.

Grasping cognitive tendencies improves the accuracy of the prototype

Once you can see your own cognitive tendencies, the accuracy of the 'cognitive prototype' mentioned in the internal training section will also increase.
It becomes possible to operate by saying things like, 'I'm prone to bias here, so I'll consciously correct it.'

Let's use external devices to continuously calibrate our prototypes 👍


Summary of the External Device Edition (throughout both parts)

  • Others are not 'mirrors,' but in many cases are just 'passersby'
    (who lack a sense of ownership)

  • Human-based external evaluation has extremely high search costs and context synchronization costs

  • The essential strength of AI is not intelligence, but the compression of question synchronization costs

  • The era of role-fixed prompts is coming to an end, and we are entering an era of co-designing evaluation functions

  • By having AI observe and record your cognitive tendencies, continuous calibration of the prototype becomes possible

By incorporating external devices, objective viewing finally begins to function properly.
Internal Training × Prototype × External Device (Evaluation Function Co-design)
Once you have all this, your objective viewing skills should be quite powerful.

……But!

In this series, there is one last, heavy topic remaining.
The biggest trap that people who have cleared all this training are most likely to fall into.

The conviction that 'I am able to view things objectively'

This is seriously troublesome.
The more one has undergone internal training, possesses a prototype, and operates at a level where they co-design evaluation functions, the more they fall into this trap.
This is because there is no mechanism yet in place to doubt the 'self that is viewing things objectively'.

Next time, finally the last installment.
Recursive Audit Edition
How do you audit the self that is viewing things objectively?

With that said, I, who am speaking so pretentiously, am the one most careful and struggling daily with this final installment of 'viewing objective perspective objectively'.

いいなと思ったら応援しよう!