Automation of AI Alignment Research and the Intelligence Explosion
I've shifted my research to focus on automated alignment research. We will have automated AI research very soon and it's important that alignment can keep up during the intelligence explosion. https://t.co/dnMt2PsPIu
— Stephen McAleer (@McaleerStephen) December 20, 2025
According to METR's evaluation, Claude Opus 4.5 has reached a 50%-time horizon (a benchmark for how long a task can be performed autonomously) of approximately 4 hours and 49 minutes in specific tasks.
This is a record high, suggesting that AI's ability to perform complex autonomous tasks (such as research and development) for long periods without human intervention is rapidly increasing.
Stephen McAleer predicts that in the very near future, AI will begin conducting AI research itself, and he has stated that he has shifted the focus of his own research to "Automated Alignment Research."
In this article, based on the importance of "Automated Alignment Research" advocated by Stephen McAleer, let's consider the grand theme of how AI should grasp "human intent" and evolve.
The General Will Perspective and Constitutional AI
As the "intelligence explosion" where AI evolves itself approaches, the most important question now is not "What should we teach AI?" but "How can we ensure AI does not deviate from humanity's true intent?".
The key to solving this difficult problem lies in the fusion of the ideas of the 18th-century philosopher Jean-Jacques Rousseau and the AI technology known as "Constitutional AI."
1. Rousseau's "General Will": What lies beyond fragmented desires
First, let's organize Rousseau's "General Will", which is the foundation of political science. He thought of people's "will" in two parts.
Particular Will (individual selfishness): Individual selfish desires, such as "I don't want to pay taxes" or "I want to take it easy."
Will of All: A simple sum of particular wills (a state close to what is known as majority rule).
General Will: A universal "correct will" that aims for the common interest of the entire society.
Rousseau called the "correct answer for all members of society," which remains after stripping away individual selfishness, the General Will. AI alignment can be said to be an attempt to implement this "General Will" into AI.
2. How Constitutional AI works
So, how do we teach this "General Will" to a digital entity like AI? A leading method for this, adopted by companies like Anthropic, is "Constitutional AI."
Conventional AI relied on "Reinforcement Learning from Human Feedback (RLHF)," where humans score each response as "this is good, this is bad." However, this mixes in human subjectivity and bias.
Mechanism of Constitutional AI:
The AI is given a "constitution" (a list of behavioral guidelines) (e.g., "do not say harmful things," "respect freedom," etc.).
The AI self-censors its own responses against that "constitution."
The AI itself judges whether it has "violated the constitution" and repeats the process to refine its responses into better ones.
In a sense, it is a mechanism that houses a "rational judge" within the AI.
3. "Human Values" in Constitutional AI
The content of the "constitution" that Constitutional AI must uphold is determined by humans. It aggregates human rights declarations and values from around the world and sets them as an ethical "North Star" that the AI must follow.
However, we face a major problem here: defining "human values" is unimaginably difficult that is.
4. "Collective Intelligence" and "Root Knowledge": A Perspective from Physics
Let us broaden our perspective a little here.
Collective Intelligence: An aggregation of many people's opinions. This is close to what Rousseau called the "general will," and it carries the risk of falling into mere majority rule.
Root Knowledge: Universal principles based on the laws of the universe and physical rationality.
Recent theories (such as the integration of physics and sociology) suggest that "ethics are rational rules for a living system to maintain energy and evolve most efficiently."
The morality of "not lying" may be a "physical optimal solution" to lower the communication costs of the social system and prevent its collapse. Alignment research is beginning to go beyond mere "human preferences" to pursue these cosmic universal principles (the ultimate form of the general will).
5. The Wall of Human Value Imperfection
We humans, who are trying to align AI, have fatal flaws.
Limits of verbalization: We cannot accurately explain what we want even to ourselves.
Contradictions: We hold contradictory values simultaneously, such as "seeking safety" while "enjoying thrills."
Bias: We cannot escape the prejudices of a specific culture or era.
If we copy the values created by "imperfect humans" directly into AI, the AI will become a "high-performance, imperfect monster."
6. Conclusion: Why We Need "Fundamental Universal Principles"
This is why Stephen McAleer is rushing toward "automated alignment."
When an "intelligence explosion" occurs where AI surpasses human intelligence, it will be too late for humans to manually correct the AI. Furthermore, with ambiguous human instructions, the AI will either run out of control or learn human stupidity.
What we must think about now is not the short-term perspective of "what humans want," but rather finding the "general will" that we should reach after intelligence continues to evolve—that is, the "physical and logical fundamental principles for the coexistence of intelligence and life."
Through the mirror of AI, we are attempting to answer the philosophical question of 'what is the human will' using the language of mathematics and physics.
