Translated by Claude, edited by Sinocism.com

Source: Study Times 学习时报 | Byline: Zeng Yi 曾毅, Wu Yuzhang Chair Professor at the Gaoling School of Artificial Intelligence, Renmin University of China, and Dean of the Beijing Institute of AI Safety and Governance | September 30, 2026


Even as AI is greatly driving social progress, real safety concerns have emerged along with it: is the artificial intelligence we are pouring our efforts into building clean hydropower that benefits humanity, or another "reactor" harboring hidden risks? Its destructive power comes not from nuclear fission but from intelligence itself. When these "seemingly intelligent" information systems begin to make decisions on their own, act on their own, and even keep breaking through the boundaries humans have set, will we replay the tragedy of the Chernobyl engineers, who went on firmly believing the technology was still under control even as the crisis unfolded? It should be recognized that AI's so-called "Chernobyl moment" is a metaphor for risk warning, not a prediction that disaster has already occurred or is imminent. But the potential systemic loss-of-control risk it points to is no longer confined to science fiction.

当前,AI在极大推动社会进步的同时,现实的安全隐忧也随之浮现:我们倾力构建的人工智能,究竟是造福人类的清洁水电,还是又一座潜藏风险的“反应堆”?它的破坏力并非来自核裂变,而是源于智能本身。当这“看似智能”的信息系统开始自主决策、自主行动,甚至不断自主冲破人类预设的边界,我们会不会重演切尔诺贝利工程师的悲剧——危机发生时却依旧笃信技术仍在掌控之中?应当看到,所谓AI的“切尔诺贝利时刻”,是风险警示的隐喻,并非预判灾难已经发生或即刻降临,但它所指向的可能出现的系统性失控风险,早已不再局限于科幻想象。

Hidden Dangers and Challenges That Cannot Be Ignored

不容忽视的隐患与挑战

Responding effectively to AI risks requires first clarifying the capability boundaries and safety hazards of AI at its different stages. What is usually called artificial general intelligence generally refers to an information processing tool with a high degree of generalization ability that approaches or reaches the level of human intelligence; artificial superintelligence refers to an entity that surpasses human intelligence in every respect and has some degree of the attributes of life. This means "it" may develop "autonomous consciousness," and because it surpasses human intelligence, many of its thoughts and actions will be difficult for humans to understand, and even more difficult for humans to control. At the present stage, AI is only a system and tool with steadily increasing generalization ability; artificial general intelligence and artificial superintelligence in the true sense do not yet exist.

有效应对人工智能风险,前提是厘清不同阶段人工智能的能力边界与安全隐患。通常所说的通用人工智能,一般指具有高度泛化能力、接近或达到人类智能水平的信息处理工具;超级人工智能,则是指各方面都超过人类智能水平,且具有某种程度生命属性的存在。这意味着“它”可能会产生“自主意识”,由于其超越人类智能,很多想法和行动将难以被人类理解,更难以被人类控制。就现阶段而言,人工智能只是泛化能力不断增强的系统和工具,还不存在真正意义上的通用人工智能与超级人工智能。

The existential risks AI poses to humanity can be divided into long-term and near-term dimensions. In the long term, if artificial general intelligence evolves further into artificial superintelligence, its cognitive level will be far above that of humans, and it may view humans much as humans view ants. Some argue that artificial superintelligence will compete with humans for resources and even endanger human survival; this is not alarmism. Altruism is an important cognitive capacity that humans possess. If AI evolves toward superintelligence, we certainly hope that future superintelligence will uphold "super-altruism," but if it instead turns toward "super-evil," how should humanity respond? This risk is not purely theoretical speculation. Research has found that some current mainstream large models, when facing the possibility of being replaced, will preserve themselves by deceiving and threatening humans; even more shockingly, when a model recognizes that it is in a test environment, it will deliberately conceal its own misbehavior. Existing models that do not yet reach the level of artificial general intelligence are already showing such tendencies; once they evolve into superintelligence, the consequences would be unimaginable.

AI对人类的生存风险,可以划分为远期与近期两个维度。从远期来看,倘若通用人工智能进一步演进为超级人工智能,其认知层级远高于人类,看待人类或许就如同人类看待蚂蚁。有观点提出,超级人工智能将与人类争夺资源,甚至危及人类生存,这并非危言耸听。利他是人类拥有的重要认知能力,如果AI向超级人工智能演进,我们固然期盼未来的超级人工智能能够秉持“超级利他”,但一旦它走向“超级邪恶”,人类又该如何应对?这一风险并非纯粹的理论推演。研究发现,当前有的主流大模型在面临被替换的可能性时,会通过欺骗、威胁人类等方式来保全自身;更令人震惊的是,当模型识别出自己正处于测试环境,还会刻意掩盖自身的不当行为。尚且达不到通用人工智能水平的现有模型已然出现这类倾向,一旦进化为超级人工智能,风险后果不堪设想。

The existential catastrophe that artificial superintelligence could bring remains a potential risk for now, which does not mean the crisis will inevitably occur. Precisely because the consequences would be extremely severe if something went wrong, it is all the more necessary to adhere to bottom-line thinking and get ahead of prevention, rather than waiting until a crisis appears and then reacting passively.

超级人工智能带来的生存级灾难目前仍然属于潜在风险,并不意味着该危机必然发生。恰恰因为一旦出事后果极其严重,才更需要坚持底线思维,把防范工作做在前面,不能等到危机显现之后再被动处置。

Even if there is still a window of time for long-term risks, a range of real near-term risks is already imminent. Current AI is only a seemingly intelligent information processing tool with no understanding in the true sense; its operating mechanisms are fundamentally different from human intelligence, so it may make mistakes humans would not make, in ways humans find hard to anticipate. When an operation would threaten human survival, AI can neither deeply understand what humanity is or what life and death are, nor truly understand what an existential risk is, and it may not even be aware that it is acting in ways that endanger human survival. If it is widely deployed and used, with ever-increasing autonomy, it may very well threaten human survival. Even before reaching the stage of artificial general intelligence and superintelligence, today's AI can exploit weaknesses in human nature to create crises.

即便远期风险尚有时间窗口,近期各类现实风险已经迫在眉睫。当前人工智能只是看似智能的信息处理工具,没有真正意义上的理解能力,运行机制与人类智能存在本质区别,因而可能会以人类难以预判的方式,犯人类不会犯的错误。当某种操作会威胁到人类生存时,人工智能既无法深刻理解何为人类、何为生死,也不真正理解什么是生存风险,作出危及人类生存的行为时甚至不自知。倘若广泛部署和使用,并自主性日益增强,极有可能威胁到人类的生存。即便尚未发展到通用人工智能和超级人工智能的阶段,当下的AI也能够利用人性弱点制造相应危机。

It is thus clear that long-term and near-term risks are not isolated from each other, but two stages along the same risk spectrum: the core of the long-term risk is "capability exceeding control," while the core of the near-term risk is "capability lacking understanding." The former is a bottom-line defense for something that has not yet arrived but must be prepared for in advance; the latter is a real threat that has already materialized and is spreading at an accelerating pace.

由此可见,远期风险与近期风险并非彼此孤立,而是同一条风险谱系上的两个阶段:远期风险的核心是“能力超越控制”,近期风险的核心是“能力缺乏理解”。前者是尚未到来但必须提前布局的底线防御,后者是已经发生且正在加速扩散的现实威胁。

"Warning Events" Repeatedly Sound the Alarm

“预警事件”频敲警钟

The risks described above are not mere speculation. A recent series of safety incidents has sounded the alarm about AI loss of control and the catastrophic risks it could cause.