What are linguistic issues in Taiwan languages that are typically misunderstood, misrepresented in state-of-the-art AI models
| Linguistic issue | Taiwan-specific manifestation | Typical AI/model failure mode | Resulting risk in Taiwan |
|---|---|---|---|
| Hokkien–Mandarin code-mixing and non-standard Han characters | Hokkien sexual or abusive expressions written with borrowed Han characters (e.g., non-standard spellings that encode explicit sexual acts) | Fails to recognize semantic equivalence to explicit Mandarin sexual or abusive vocabulary; classifies as benign or weakly sensitive | Unflagged sexual and abusive content, especially harmful to minors and to services that rely on lexical-level filters |
| Traditional Chinese mixed with Zhuyin and informal symbols | Use of 注音符號 (Zhuyin), emoji, and stylized text interspersed with Traditional Chinese characters in everyday writing | Cannot robustly normalize or interpret mixed-script strings; harmful |