问HN:你还能在自己的语言中区分出AI生成的文本吗?
我是日本人,通常可以通过用词来识别AI生成的日语。有些词,比如“実務”(实际工作)和“帳簿”(账本),出现的频率远高于正常情况。像“効く”(有效)和“刺さる”(打动人心)这样的随意用词,通常会在本该使用其他词的地方出现。它们就像英语中的“delve”。
我并不是以英语为母语的人,但我觉得AI生成的英语最近有了很大改善,比以前更难以识别。这和英语母语者的感受一致吗?
其他语言呢?中文一定是训练得最充分的语言之一,所以我期待它的改善程度与英语相当。那么西班牙语、印地语以及其他广泛使用的语言呢——说话者还能分辨出来吗?
此外,我认为一旦离开英语,所使用的模型仍然非常重要。我第一次感受到这一点是在Gmail的AI草拟功能中。它显然吸收了日本企业邮件文化的惯例,包括官僚主义。
如今,Anthropic的Opus 4.6(以及在编码会话外的Fable 5)能够写出相当不错的长篇日语,而谷歌的模型在更广泛的语域中表现良好。OpenAI的模型,Opus 4.8/5和在编码会话中的Fable听起来过于书呆子气。
期待听到你的想法。
查看原文
I'm Japanese, and I can <i>still</i> usually spot AI-generated Japanese by its word choice. Some words, like 実務 (practical work) and 帳簿 (a ledger), turn up far more often than they should. Casual ones like 効く (to work well, to be effective) and 刺さる (to pierce > to resonate, to hit home) get used where another word would normally be the first choice. They're the Japanese equivalent of "delve."<p>I'm not a native English speaker, but my impression is that AI-generated English has improved a lot recently and is harder to pick out than it used to be. Does that match what English native speakers see?<p>What about other languages? Chinese must be one of the most heavily trained languages, so I'd expect it to have improved as much as English. And what about Spanish, Hindi, and other widely spoken languages — can speakers still tell?<p>Also, I think which model you use still matters far more once you leave English. One of my first singularity moments came from Gmail's AI drafting feature. It had clearly absorbed the conventions of Japanese corporate email culture, bureaucracy and all.<p>These days, Anthropic's Opus 4.6 (and Fable 5 outside coding sessions) writes reasonably good long-form Japanese, and Google's models hold up across a wider range of registers. OpenAI's models, Opus 4.8/5 and Fable in a coding session sound way too geeky.<p>Curious to hear your thoughts.