氯雷他定 (@ziting_liu) 在 关于Anthropic 最新的 AI 滥用报告总结 中发帖
这篇报告是 Anthropic 于 2026 年 9 月发布的《Detecting and countering misuse of AI: September 2026》,涵盖 2025 年 12 月至 2026 年 8 月期间其威胁情报团队发现并处置的滥用案例。报告涉及中国大模型公司的部分集中在 “Illicit distillation(非法蒸馏)” 章节,同时也散见于 Cyber operations 和 Surveillance operations 中与中国相关的国家支持型滥用案例。以下按主题梳理涉及国内大模型的事情及质控问题。
一、涉及国内大模型公司的非法蒸馏(Illicit Distillation)
报告明确指出,自 2026 年 2 月首次披露以来,Anthropic 已识别并 disrupt 了来自 七家中国实验室 的蒸馏攻击,全部针对其一般可用模型(未涉及 My...