请问HN:人工智能水印是否带来了新的攻击向量?
我从Claude关于水印的文档中尚未确定的一件事是,他们在与特定水印关联时存储了什么类型的元数据。显然,他们可以使指纹尽可能独特,甚至精确到具体的时间、用户和会话。
如果是这样,这似乎是一个AI用户可能没有考虑到的隐性风险。您现在编写的任何代码都携带着您可能不希望被揭示的信息。如果恶意行为者获得了密钥,他们可能能够揭露那些希望保持匿名的开源贡献者。或者,跨样本的足够多的指纹可能会揭示出公司不愿意公开的内部组织细节。
拥有更多想象力的人可能会想出更好的例子,这似乎是一个我没有看到太多关注的攻击向量。
查看原文
One thing I have not been able to determine from Claude's documentation on watermarking[0] is what kind of metadata they store in association with a given watermark. Ostensibly they could make the fingerprints as unique as they want, possibly down to the exact time, user and session.<p>If so, this seems like hidden risk that AI users are probably not considering. Any code you write now carries information that you might not want revealed. If a bad actor gets the keys then they may be able to de-anonymous open source contributors who want to stay hidden. Or perhaps enough fingerprints across a sample could reveal internal organization details a company would rather not disclose.<p>Someone with more imagination can probably come up with better examples, it just seems like an attack vector that I haven't seen much consideration for.<p>[0] https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content