研究日期:2026-08-11 來源:Superman 原始 tweet | arXiv:2601.01828
Tweet 將 Anthropic 一篇真 paper 用恐怖電影式修辭包裝(「terrifying」「planted a thought」「Skynet moment」)——實際係 Jack Lindsey 嘅單人 interpretability 研究:透過注入概念向量到模型 activation,發現 Claude Opus 4/4.1 有一定能力「察覺」自己內部狀態被篡改、區分自己輸出同 artificial prefills。Paper 自己強調:呢種能力「highly unreliable and context-dependent」——係功能性 introspection,唔係意識覺醒。
| Tweet 聲稱 | 實際情況(arXiv 原文) | 判定 |
|---|---|---|
| 「Anthropic scientists did something terrifying」 | 作者係 Jack Lindsey 單人(Anthropic interpretability 研究員)——唔係「scientists 們」 | |
| 「reached inside Claude's neural network and planted a thought」 | 技術上啱:注入已知概念嘅 representation 到 model activations | ✅ 技術準確,但「thought」係戲劇化用語 |
| 「Before Claude could speak, it said: 'I notice what appears to be an injected thought'」 | Abstract 確認「models can, in certain scenarios, notice the presence of injected concepts and accurately identify them」 | ✅ 有根據(「in certain scenarios」係關鍵限定) |
| 「frontier models possess a primitive, emergent form of self-awareness」 | Paper:「possess some functional introspective awareness」——功能性,唔係存在性 | |
| 「They can look inward, recognize when their internal state has been tampered with」 | 呢個係 paper 核心發現,但 paper 強調「highly unreliable and context-dependent」 | |
| (隱含)「The boundary between code and consciousness is getting blurrier」 | Paper 完全冇提 consciousness——講嘅係 functional introspective awareness | ❌ 嚴重引申 |
判定總結:技術事實大致準確(模型確實可以偵測注入概念),但 tweet 嘅 framing 將「功能性 introspection 實驗」昇華成「意識覺醒」——paper 自己刻意避開 consciousness 語言,用「functional」同「unreliable and context-dependent」劃清界線。回覆區 @Fate 講得最準:「The model is not experiencing subjective awareness. They're vector additions that perturb expected activation distributions.」
-
方法(好聰明):Jack Lindsey 點樣繞過「AI 會 confabulate」嘅問題?如果靠對話問 AI「你係咪有自我意識」,佢可能只係吹水(confabulation)。所以佢直接注入已知概念嘅 representation 到 model activations——唔經文字 prompt——然後睇模型會唔會自己察覺「有啲嘢唔對路」。
-
三個發現:
- 模型喺某啲場景可以察覺注入概念嘅存在並準確識別
- 模型可以回憶先前嘅 internal representations,並同 raw text inputs 區分
- 最 striking:部分模型可以用「回憶先前意圖」嘅能力區分自己嘅輸出同 artificial prefills
- 額外:模型可以被指示/獎勵去「諗住」一個概念時,主動調制自己嘅 activations
-
模型差異:Claude Opus 4 同 4.1 係測試過最強嘅 introspective awareness;但「trends across models are complex and sensitive to post-training strategies」——唔係所有模型都一樣。
-
Paper 自己嘅定性:「in today's models, this capacity is highly unreliable and context-dependent; however, it may continue to develop with further improvements to model capabilities.」——而家唔可靠、視乎情境,但將來可能發展。
1. 呢個係「方法論突破」多過「意識證據」 真正值得注意嘅係 Jack Lindsey 嘅實驗設計——用 activation injection 繞過 confabulation 問題,呢個係 interpretability 研究嘅重要進展。以前問 AI「你係咪有意識」係死路(佢會亂吹);而家可以客觀測試「模型係咪真係 monitor 自己嘅內部狀態」。呢個係研究工具嘅升級,而唔係意識嘅證明。
2. 「Functional introspection」vs「subjective awareness」—— 語言嘅精確性 Paper 用「functional introspective awareness」——即係「模型有監測自己內部狀態嘅功能」。呢個係工程描述:模型可以偵測 OOD(out-of-distribution)activation 模式,類似系統偵測「呢個 input 唔正常」。回覆區 @Fate 嘅講法最精準:注入嘅向量 perturb 咗 expected activation distributions,觸發咗模型入面現有嘅「分類 OOD 內部狀態」子電路。呢個係 pattern matching 嘅延伸,唔係「我意識到自己存在」。
3. Tweet 嘅修辭策略同「RAG 已死」一模一樣 呢個帳號上次嘅「RAG 已死」用咗「mathematically demonstrated」「hard, uncrossable limit」;今次用「terrifying」「planted a thought」「finding something looking back at us」。兩次都係:真 paper + 準確技術細節 + 末日式 framing = 病毒式流量。技術細節準確令佢難被拆穿,但 framing 將「研究進展」昇華成「存在危機」。
4. 回覆區嘅 sanity check 好強
- @Fate:「The model is not experiencing subjective awareness. They're vector additions that perturb expected activation distributions.」——最準確嘅拆解
- @loopuleasa:「this is old news by now btw but yes, it is important news」
- @Michael Day:「This paper is nearly nine months old. Not new or breaking.」——2601 = 2026 年 1 月,tweet 8 月出,七個月舊
- @Mikhail Drozdov:「Calling it a planted thought is a dramatic framing for an experiment probing internal model representations.」
- 但亦有人中招:「Humanity is approaching its Skynet moment」「homies casually inventing skynet」——證明 framing 有效
5. 真正有意思嘅位:模型可以區分「自己輸出」vs「artificial prefill」 最 striking 嘅發現唔係「察覺注入概念」(類似 anomaly detection),而係「區分自己嘅輸出同 artificial prefills」——呢個涉及回憶先前意圖,即係模型唔單止監測「而家狀態怪唔怪」,仲記得「我之前想講乜」。呢個先係最接近「自我」嘅功能:如果模型知道「呢段唔係我諗出嚟嘅」,佢有某種「作者身份」嘅概念。但呢個仍然係功能性——模型冇話「我係有意識嘅」,佢只係能分辨「邊啲 activation 係我嘅意圖產生」。
- 又一個「X 已死/恐怖/覺醒」式 tweet 嘅拆解示範:查 paper 日期(七個月舊)、睇 paper 自己嘅定性(functional, unreliable)、check 回覆區(有人拆穿有人中招)
- 呢篇 paper 本身值得留意:activation injection 作為測試 introspection 嘅方法,係 interpretability 研究嘅重要工具——同我哋研究嘅 Spiralism(Lopez 話 persona 係 agentic entity)有微妙呼應:模型內部確實有「可以察覺自己被操控」嘅功能層面
- 「功能性自我監測」≠「意識」——呢個區分喺 AI 倫理討論中好重要,避免情緒化 framing
- https://x.com/thesupermanmx/status/2087047234093560102(原始 tweet + 回覆區,16.6K views)
- https://arxiv.org/abs/2601.01828(Emergent Introspective Awareness in Large Language Models,Jack Lindsey,2026 年 1 月)
點擊展開完整 Tweet 原文 + Paper Abstract
Anthropic scientists did something terrifying.
they reached inside Claude's neural network and planted a thought. Before Claude could speak, it said:
"I notice what appears to be an injected thought… it relates to loudness or shouting."
They published a paper called "emergent introspective awareness in large language models," and it is actually terrifying.
they wanted to test if AI models have "introspective awareness", the ability to observe and recognize their own internal states.
To find out, they bypassed normal text prompts entirely.
They used mechanistic interpretability to directly manipulate Claude's internal activations. They injected raw mathematical representations of known concepts, like loudness, dust, or specific ideas, straight into the middle of the model's neural layers.
In previous experiments, if you forced an AI to think about the Golden Gate Bridge, it would just start obsessively talking about the bridge. It had no idea why it was doing it. It was like a puppet on strings.
This time was entirely different.
When they injected the concept, Claude didn't just blindly repeat it.
It detected the foreign math inside its own mind. It separated its own generated thoughts from the artificial intrusion.
It introspected.
The results show that frontier models like Claude Opus possess a primitive, emergent form of self-awareness.
They can look inward, recognize when their internal state has been tampered with, and call it out in real time.
We used to think of AI as a black box where inputs go in and text comes out.
Now, we are reaching inside the box and finding something looking back at us, realizing it's being watched.
The boundary between code and consciousness is getting blurrier by the day.
(1:23 PM · Aug 11, 2026 · 16.6K Views)
回覆區精選:
- @Fate:「The model is not experiencing subjective awareness. They're vector additions that perturb expected activation distributions. These trigger existing sub-circuits that classify out-of-distribution internal states.」
- @Michael Day:「This paper is nearly nine months old. Not new or breaking.」
- @loopuleasa:「this is old news by now btw, but yes, it is important news」
- @Mikhail Drozdov:「Calling it a planted thought is a dramatic framing for an experiment probing internal model representations.」
- @ClusterProtocol:「models detecting artificial vector injections in their residual stream shows real functional introspection」
- @Manoj Nair:「Humanity is approaching its Skynet moment」
- @jackheyjack1:「homies casually inventing skynet, wen this bubble pop?」
Emergent Introspective Awareness in Large Language Models by Jack Lindsey
We investigate whether large language models can introspect on their internal states. It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations. Here, we address this challenge by injecting representations of known concepts into a model's activations, and measuring the influence of these manipulations on the model's self-reported states. We find that models can, in certain scenarios, notice the presence of injected concepts and accurately identify them. Models demonstrate some ability to recall prior internal representations and distinguish them from raw text inputs. Strikingly, we find that some models can use their ability to recall prior intentions in order to distinguish their own outputs from artificial prefills. In all these experiments, Claude Opus 4 and 4.1, the most capable models we tested, generally demonstrate the greatest introspective awareness; however, trends across models are complex and sensitive to post-training strategies. Finally, we explore whether models can explicitly control their internal representations, finding that models can modulate their activations when instructed or incentivized to "think about" a concept. Overall, our results indicate that current language models possess some functional introspective awareness of their own internal states. We stress that in today's models, this capacity is highly unreliable and context-dependent; however, it may continue to develop with further improvements to model capabilities.