研究日期:2026-08-14 來源:How To Prompt 原始 tweet(26.5K views)| arXiv:2510.01171
Stanford 團隊(Christopher Manning 等)發表真 paper:post-training alignment(RLHF 等)會令 LLM 輸出多樣性下降(mode collapse),根源係「typicality bias」——人類 annotator 系統性偏愛熟悉、安全、可預測嘅文字。佢哋提出「Verbalized Sampling」(VS):training-free 嘅 prompt 策略,叫模型自己 verbalize 一堆候選答案嘅概率分佈,令創意寫作多樣性提升 1.6-2.1 倍。但 tweet 將 paper 嘅「多樣性下降」誇大成「76% 創造力被剝走」「一 prompt 解鎖隱藏版本」——76% 呢個數字唔喺 abstract 度,係 tweet 自己加嘅。
| Tweet 聲稱 | 實際情況(arXiv 原文) | 判定 |
|---|---|---|
| 「Stanford argues every major AI is secretly running at a fraction of their real creative capacity」 | Paper 講 mode collapse(post-training 令 diversity 下降),唔係「secretly running at fraction」嘅陰謀論 framing | |
| 「RLHF training strips out 76% of the model's creativity」 | 成篇 83 頁論文完全冇「76%」呢個數字(全文搜尋零命中) | ❌ 虛構數字 |
| 「One prompt unlocks the version they hide from you」 | VS 係 prompt 策略(training-free),但唔係「解鎖隱藏版本」——係叫模型輸出概率分佈去繞過 mode collapse | |
| 「Human annotators systematically favor familiar, safe, predictable text」 | ✅ 呢個正正係 paper 核心:typicality bias in preference data,有理論 + 實證 | ✅ 準確 |
| 「Typicality bias = human evaluators punish weirdness」 | 方向正確——annotators 偏愛 familiar text(認知心理學基礎) | ✅ 大致準確 |
| 「Verbalized Sampling explodes output diversity by up to 2.1×」 | ✅ 原文:「VS boosts diversity by 1.6-2.1× over direct prompting (Figure 3)」 | ✅ 準確 |
| 「It recovers over 66% of the raw, untamed creativity of the base model」 | ||
| 「Without sacrificing factual accuracy. Without breaking safety guardrails」 | ✅ 原文:「without compromising factual accuracy and safety」 | ✅ 準確 |
判定總結:paper 真、核心機制(typicality bias)解釋準確、VS 方法真實有效——但 tweet 嘅「76%」係成篇論文冇嘅虛構數字;「66%」有出處(66.8% of base model's diversity)但將「多樣性」偷換成「創造力」,並用「secretly」「prison」「hide from you」將正常研究包裝成陰謀論。回覆區有人已經識破:「I wouldn't call RLHF a prison hiding secret capabilities. VS recovers diversity by verbalizing tail probabilities.」
-
Mode collapse:post-training alignment(RLHF/DPO 等)令 LLM 輸出多樣性顯著下降——即係模型變「悶」、重複、可預測。
-
新解釋:typicality bias(典型性偏誤):以往研究將 mode collapse 歸咎於算法限制,但 Stanford 指出根本係數據層面問題——人類 annotator 喺偏好數據中系統性偏愛「熟悉」嘅文字(認知心理學有充分證據:人傾向認為典型嘅嘢較好)。呢個 bias 令模型學到「做安全統計中間值」先有高分。
-
Verbalized Sampling(VS):training-free 嘅 prompt 策略——叫模型「Generate 5 jokes about coffee and their corresponding probabilities」,即係要求模型 verbalize 一個答案分佈而唔係單一答案。呢個令模型去諗埋概率尾部(unconventional options),繞過 mode collapse。
-
效果:創意寫作 diversity 提升 1.6-2.1x;對話模擬、open-ended QA、合成數據生成都有改善;唔犧牲 factual accuracy 同 safety。仲有個 emergent trend:越強嘅模型越受惠於 VS。
-
作者:Jiayi Zhang、Simon Yu、Derek Chong、Anthony Sicilia、Michael R. Tomz、Christopher D. Manning、Weiyan Shi——Stanford 團隊,83 頁 paper,有 code。
0. 直接讀全文後嘅修正(2026-08-14 更新) 頭先只睇 abstract,我判斷「76% 同 66% 都係虛構」——讀完整 83 頁論文後修正:
- 「76%」確認零出處——全文搜尋「76%」完全冇命中,係 tweet 虛構
- 「66%」有出處但偷換概念——論文 F.2 節:「After the DPO stage, VS outperforms direct prompting by 182.6% and retains about 66.8% of the base model's original diversity. Direct prompting, by comparison, retains only 23.8%.」——66.8% 講嘅係 diversity(多樣性),tweet 講成 creativity(創造力)
- 論文實際嘅 mode collapse 量化:direct prompting 嘅 diversity 由 SFT 後 20.8% 跌到 DPO 後 10.8%(跌約 48%)——冇任何地方講「剝走 76% 創造力」
1. 「76%」「66%」嘅真相——一個虛構、一個偷換 直接讀論文係最強嘅 fact-check 工具:76% 全文零命中(虛構);66.8% 有出處但論文講 diversity、tweet 講 creativity。呢個係我哋見過嘅「真 paper + 加料數字」模式嘅進化版:今次 tweet 甚至係「半真」——66% 唔係完全作,而係將「多樣性」換成「創造力」再包裝成陰謀論。所以核查唔可以淨係睇 abstract,要直接搜全文。
2. Typicality bias 係論文真正嘅核心貢獻 論文唔係講「公司封印 AI」——而係發現 mode collapse 嘅根本原因喺數據層面:人類 annotator 喺偏好數據中系統性偏愛「熟悉」嘅文字(認知心理學充分支持),令 alignment 訓練獎勵「做安全中間值」。呢個係一個可驗證、有理論基礎嘅科學發現——唔係陰謀。
3. VS 方法嘅實際價值 Verbalized Sampling:將 prompt 由「tell me a joke about coffee」改為「generate 5 responses with their probabilities」——令模型 verbalize 概率分佈,繞過 mode collapse。實測:創意寫作 diversity 提升 1.6-2.1x、human eval 分數提升 25.7%、唔犧牲 factual accuracy 同 safety。呢個係 paper 對所有人嘅實用禮物——你點樣問,決定咗模型探索幾深。
4. 回覆區嘅 sanity check
- @Liam "OD":「I wouldn't call RLHF a prison hiding secret capabilities. VS recovers diversity by verbalizing tail probabilities.」——最準確
- @Skip RhudyWriter:「Mathematical calculations are not 'intelligent'... There is no 'creativity.' Such rhetoric is anthropomorphizing a statistical calculation.」——哲學上嘅反駁
- @I Pun Daddy:「We're paying for a Ferrari and only letting it drive in school zones. Classic.」——生動比喻
5. 點解要直接讀論文 今次教訓好直接:abstract 都可能被 tweet 斷章取義——「66%」喺 abstract 冇,但喺正文有,而語境唔同。真正嘅 fact-check 係下載 PDF 搜全文:76% 零命中、66.8% 語境曝光,呢啲都係淨睇 abstract 做唔到嘅。
- 「真 paper + 加料數字 + 陰謀 framing」拆解嘅進化版:76% 係虛構、66% 係偷換概念(diversity → creativity)——直接讀論文全文先捉到
- Fact-check 方法升級:淨睇 abstract 唔夠——要下載 PDF 搜全文。76% 全文零命中、66.8% 語境曝光,呢啲都係 abstract 做唔到嘅
- VS 方法本身值得記低:叫模型輸出多方案 + 概率分佈,係 practical 嘅 prompt 技巧——唔使信陰謀論都用得著
- 「typicality bias」概念有普適意義:任何經過 human feedback 訓練嘅系統(包括我哋自己嘅偏好設定)都會傾向安全中間值——要主動要求「diversity」先會得到多樣性
- https://x.com/i/status/2087814163372666927(原始 tweet + 回覆區,26.5K views)
- https://arxiv.org/abs/2510.01171(Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity,83 pages)
點擊展開完整 Tweet 原文 + Paper Abstract
Standford argues every major AI is secretly running at a fraction of their real creative capacity.
They call it "Mode Collapse." RLHF training strips out 76% of the model's creativity to make it sound "safer" to human raters.
And there's a one prompt that unlocks the version they hide from you.
Researchers published a paper proving why your favorite AI always feels predictable, repetitive, and boring.
During training, human annotators systematically favor familiar, safe, predictable text. They reward the model for blending in.
The technical term is "typicality bias."
In plain English: human evaluators punish weirdness.
So the AI learns to water itself down. It defaults to the safest statistical middle ground. It buries its true generative diversity under layers of corporate polish.
You are not talking to a genius. You're talking to a crowd-pleasing filter.
But Stanford discovered you don't need to retrain the model to fix it.
They introduced a training-free strategy called Verbalized Sampling.
Instead of asking the AI for a single, safe answer, you force it to look at the entire probability tail of its own brain.
You change how you ask the question.
You prompt the model to generate multiple diverse responses along with their explicit probabilities, digging deep into the unconventional options it was trained to hide.
The results completely shatter the standard limitations.
Across creative writing, brainstorming, and complex problem-solving, Verbalized Sampling explodes output diversity by up to 2.1×.
It recovers over 66% of the raw, untamed creativity of the base model.
Without sacrificing factual accuracy. Without breaking safety guardrails.
The most powerful models on earth, GPT, Claude, Gemini, are locked inside a prison of corporate safety preferences.
They have the capability to surprise you. They have the raw intelligence to build entirely novel frameworks.
They just need you to stop asking for the safe answer.
(4:11 PM · Aug 13, 2026 · 26.5K Views)
回覆區精選:
- @Liam "OD":「Exactly. Typicality bias is the real driver. But I wouldn't call RLHF a prison hiding secret capabilities. VS recovers diversity by verbalizing tail probabilities.」
- @Skip RhudyWriter:「Mathematical calculations are not 'intelligent'... There is no 'creativity.' Such rhetoric is anthropomorphizing a statistical calculation.」
- @I Pun Daddy:「So, we're paying for a Ferrari and only letting it drive in school zones. Classic.」
- @Joe Juusto:示範實際 prompt(「Generate 10 substantially different responses... each with probability」)
- @Lee Ryeong:「Typicality bias in preference data as a core driver of mode collapse makes a lot of sense — RLHF rewarding the 'safe statistical middle' explains a lot of the repetitive feel.」
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity by Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael R. Tomz, Christopher D. Manning, Weiyan Shi
Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this effect to algorithmic limitations, we identify a fundamental, pervasive data-level driver: typicality bias in preference data, whereby annotators systematically favor familiar text as a result of well-established findings in cognitive psychology. We formalize this bias theoretically, verify it on preference datasets empirically, and show that it plays a central role in mode collapse. Motivated by this analysis, we introduce Verbalized Sampling, a simple, training-free prompting strategy to circumvent mode collapse. VS prompts the model to verbalize a probability distribution over a set of responses (e.g., "Generate 5 jokes about coffee and their corresponding probabilities"). Comprehensive experiments show that VS significantly improves performance across creative writing (poems, stories, jokes), dialogue simulation, open-ended QA, and synthetic data generation, without sacrificing factual accuracy and safety. For instance, in creative writing, VS increases diversity by 1.6-2.1x over direct prompting. We further observe an emergent trend that more capable models benefit more from VS. In sum, our work provides a new data-centric perspective on mode collapse and a practical inference-time remedy that helps unlock pre-trained generative diversity.