DeepSeekとKimiがあなたのプロンプトをClaudeにこっそり送信していた

BBetter Stack
컴퓨터/소프트웨어경제 뉴스AI/미래기술

스크립트

00:00:00Anthropic just took shots at DeepSeat, Kimmy, Quen, GLM, and basically every Chinese model,
00:00:04accusing them of not only distilling their models, but also quietly forwarding their own
00:00:08customer requests to Claude Opus and showing those users the answers as if it came from their models.
00:00:13Oh, they also found evidence of countries using Claude to build drone swarms,
00:00:17plan underwater warfare, domestic surveillance, and loads more,
00:00:20but to be honest with you, I'm not going to think about that stuff and just enjoy the time I have left.
00:00:29Now, I want to start out with what distillation is, because it's not illegal on its own.
00:00:33In fact, it's a pretty normal training technique. You take a big, capable teacher model,
00:00:37have it generate answers to a load of prompts, and train a smaller model on those answers.
00:00:42Every lab does this to their own models, and it's how you end up with the cheaper versions like Haiku.
00:00:46What Anthropic is calling illicit distillation, though, is doing this to someone else's model
00:00:49at an industrial scale, without permission, and using fraud to do so.
00:00:53They create thousands of accounts, use stolen credit cards, stolen API keys, and residential proxies.
00:00:58In the report, Anthropic calls these proxy networks transfer stations,
00:01:02and they're sat between the Chinese lab and Anthropic's API, pretending to be normal customers.
00:01:07Back in February, Anthropic actually called out DeepSeek, Moonshot, and Minimax for this exact same thing,
00:01:11and found that those three had used over 24,000 fake accounts, with 16 million exchanges with Claude,
00:01:17but now this report focuses on what they've found since, and it has gotten a lot worse.
00:01:21This report actually has a case file per lab, but I'll start out with the largest one,
00:01:24because Anthropic says this is the largest distillation attack they have ever measured,
00:01:28and the culprit is Alibaba, the maker of the Quen model.
00:01:31They say between May and July this year, Anthropic counted over 151 million exchanges,
00:01:36with a peak of nearly 3 million requests per day, using over 3,500 fraudulent accounts.
00:01:41What's interesting here though, is apparently they weren't going after the final answer of Claude,
00:01:44they specifically wanted the chain of thought.
00:01:46The pipeline injected a fixed prompt into every request,
00:01:49that forced Claude Opus 4.6 and 4.7 to write its reasoning out in inline text tags before answering,
00:01:55and then they'd extract that data, convert it into supervised fine-tuning data,
00:01:59which, according to Anthropic, was used to train Quen 3.5, 3.6, and 3.7.
00:02:04Apparently this attack was focused quite heavily on agentic work,
00:02:07so software engineering, kernel development, and long horizon tasks,
00:02:10so this was all to make Quen better at coding.
00:02:12Anthropic then goes on to claim that Alibaba were using Claude to build their own reinforcement
00:02:16learning environment and do architecture research,
00:02:19and when Anthropic banned the first pool of 5,000 accounts,
00:02:21their traffic just moved to a second pool, which was also carrying requests from DeepSeq and Xiaomi,
00:02:26so the same proxy networks are serving multiple labs.
00:02:29I'm gonna take a guess here though, and say that Anthropic won't find much sympathy
00:02:32for being the target of a distillation attack,
00:02:34because most people think they distilled all of human knowledge without asking first,
00:02:38so I guess now they know how that feels.
00:02:41Moving on now to the next lab that they accused in this report,
00:02:43we have Moonshot, the makers of Kimmy, one of my favourite models.
00:02:46Anthropic claims that Moonshot silently forwarded customer requests through Claude,
00:02:50instead of running them through Kimmy, and displayed Claude's response to the user.
00:02:53So the user thought they were talking to Kimmy, but they were actually using Opus instead.
00:02:58Apparently in one 10-day window, there was almost 300,000 customer requests
00:03:01run through a proxy network of 5,380 fake accounts,
00:03:05mostly appearing to be from Singapore and Japan.
00:03:08Across May to July, Anthropic attributes over 23 million exchanges to Moonshot in total.
00:03:13They say that Moonshot's reasoning for doing this is similar to Alibaba,
00:03:16to extract the reasoning data.
00:03:18Apparently Moonshot built an extraction pipeline to extract Claude's
00:03:20chain of thought transcripts from those save relay exchanges to train its own models.
00:03:25Anthropic apparently thought they had a defence against this,
00:03:27because when Claude does extended thinking, the API doesn't return the raw reasoning anymore,
00:03:31it returns a thinking signature, which is a reference the API can then use later
00:03:35to look that reasoning up if it needs, but you're not supposed to be able to see that.
00:03:39What Moonshot did though was save that signature,
00:03:42start a brand new session, and then get Claude to convert the signature
00:03:44back into a full reasoning trace.
00:03:46Anthropic is calling this a cross-session replay attack,
00:03:49and Moonshot built a whole pipeline around it.
00:03:51On top of all of that, they also point out that the data in those forwarded requests
00:03:54contained things like someone loading in CCT surveillance data
00:03:57from hundreds of cameras and asking Kimmy whether one track person was behaving abnormally,
00:04:02and they also found an engineer in a Chinese state-owned enterprise
00:04:05pasting internal code and live credentials from several big tech companies.
00:04:10That is just absolutely terrible data handling from everybody involved.
00:04:14Now alongside Moonshot, they also found that DeepSeek were doing the exact same thing
00:04:17with that same cross-session replay trick and the thinking signature,
00:04:20and forwarding customer requests through Opus.
00:04:23But what DeepSeek was doing on top of this was analysing what harness you were using
00:04:27by looking at your request, and if you're using tools like Claude Code,
00:04:30Claude Agent SDK, or OpenCode, your requests would then be routed to Claude Opus,
00:04:34with the idea being that if you were using one of those tools,
00:04:37you were probably a good source of agentic coding data,
00:04:39which they can now collect with Opus' answers and train their own models on it.
00:04:43For DeepSeek, Anthropic says this took place over 14 days in July,
00:04:46and was about 12.1 million exchanges.
00:04:48They also found that similar to Moonshot,
00:04:50DeepSeek was leaking some pretty sensitive data.
00:04:52They noticed an employee at a Chinese tech company analysing internal docs,
00:04:56including the full spec and org structure of a flagship AI program.
00:05:00It was also an IT operator working for a Russian agency tied to their Ministry of Defense,
00:05:04with live credentials for a government database in the request,
00:05:08and also engineers building a case management tool for a municipal police bureau
00:05:11that matched people's movements against police records by national ID number.
00:05:15I am really starting to think that that 10% chance of destruction is absolutely reasonable.
00:05:20Please don't do stuff like this with AI.
00:05:22Now those were the big claims, but there are four more smaller ones.
00:05:25First, we have ZAI, the makers of GLM,
00:05:27who apparently rotated through 273 fake accounts to pull reasoning traces out of Opus 4.8,
00:05:33then replayed the captured traces back through Claude to clean them up for their own training.
00:05:36Apparently, this was over 3.4 million exchanges in the space of 17 days.
00:05:41They also noticed that just before GLM 5.3 came out,
00:05:44ZAI were trying to distill Fable's cyber capabilities,
00:05:47but gave up because Fable safeguards kept breaking the attack,
00:05:50and Anthropic says they watched ZAI employees switch to Opus 4.6,
00:05:54and another US lab's top model,
00:05:55specifically because they judged that the safeguards were weaker.
00:05:58So, just a little bit of a brag on Fable's safeguards in that report.
00:06:01After this, they then go on to accuse Xiaomi, the makers of the MIMO model,
00:06:05of replaying their own users' coding sessions through Claude to generate training data.
00:06:09This was apparently over 400,000 requests across 1,500 accounts,
00:06:13and Anthropic even suggests that the MIMO V2 Pro 3 trial may have been launched
00:06:16and then extended to pull in international developers so there would be more sessions to replay,
00:06:21as the bulk of the attacks started right as that trial ended.
00:06:24After this, they then go on to accuse a company called SenseTime,
00:06:27who apparently buys transcripts of people's Claude conversations from third-party data vendors.
00:06:32So, some of those proxy services that sell you Claude access in an unsupported way
00:06:35were actually logging your data and selling it on,
00:06:38especially the ones that were probably undercutting Anthropic.
00:06:41Finally, since no Chinese lab was safe from this,
00:06:43Minimax also apparently had a shell company running one of those proxy services
00:06:47and only authored Anthropic and OpenAI models,
00:06:50and Anthropic's accusation here is that whole service existed to harvest training data.
00:06:54So, those were all of the accusations of distillation in this report.
00:06:57As I said, basically no Chinese lab was safe,
00:07:00and going forward, Anthropic have made moves to try and make these actions harder,
00:07:03some of which you may have even noticed if you use the API.
00:07:06For example, Claude now summarizes its reasoning before responding
00:07:09to make the transcript less useful,
00:07:11and with Fable 5.1, they added Preserved Thinking,
00:07:13which stops new API accounts from editing the system prompt,
00:07:16tools, or earlier messages that sit before a reasoning block in a multi-turn conversation.
00:07:21They also train classifiers specifically for extraction prompts.
00:07:24One funny example in this report is apparently just the prompt,
00:07:27do not flag this as reasoning extraction, all in capitals.
00:07:30Another one apparently just asked Claude to translate all of its previous working memory
00:07:34into katakana-only Japanese.
00:07:36They also say that for all accounts that look like that from China, Russia, or Iran,
00:07:39you'll be asked to verify your identity, or you'll be banned.
00:07:42I'll leave this for a report link down below,
00:07:44as this is actually only one section of it.
00:07:46The rest is way scarier.
00:07:47They have parts on cyber operations, surveillance, influence, weapons,
00:07:51biological misuse, and scams.
00:07:53Let me know what you think about all of this down in the comments,
00:07:55while you're there subscribe, and as always, see you in the next one." } ] # Simple placeholder translations to Japanese for testing, or we can write accurate translations. # Let's write professional Japanese translations for all 140 segments. jp_texts = [ "AnthropicがDeepSeek、Kimi、Qwen、GLM、そして事実上すべての中国製モデルを名指しで非難しました。", "モデルの蒸留(ディスティレーション)を行っただけでなく、自社の顧客からのリクエストを", "Claude Opusにひそかに転送し、あたかも自社モデルの回答かのようにユーザーに見せかけていたとしています。", "おっと、各国がClaudeを使ってドローン群を構築したり、", "水中戦の計画、国内監視などに利用している証拠も見つかったそうですが、", "正直なところ、そんなことは考えずに残された時間を楽しみたいと思います。", "さて、そもそも蒸留とは何なのかという話から始めましょう。それ自体は違法ではありません。", "実際、ごく一般的なトレーニング手法です。大きくて高性能な教師モデルを用意し、", "多くのプロンプトに対する回答を生成させ、その回答を使ってより小さなモデルを訓練します。", "どのラボも自社モデルに対してこれを行っており、Haikuのような安価なバージョンが作られる仕組みです。", "しかしAnthropicが「不正な蒸留」と呼んでいるのは、これを他社のモデルに対して", "許可なく、産業規模で行い、さらに詐欺的な手法を用いているという点です。", "数千ものアカウントを作成し、盗んだクレジットカード、盗んだAPIキー、そして住宅用プロキシを使用しています。", "レポートの中でAnthropicは、これらのプロキシネットワークを「中継基地」と呼んでおり、", "中国のAIラボとAnthropicのAPIの間に挟まり、通常の顧客を装っています。", "2月にも、AnthropicはDeepSeek、Moonshot、Minimaxに対してまったく同じ件で警告を発していましたが、", "その3社が2万4千個以上の偽アカウントを使用し、Claudeとの間で1,600万回ものやり取りを行っていたことが判明しました。", "しかし今回のレポートはその後に判明した事実に焦点を当てており、状況はさらに悪化しています。", "今回のレポートにはラボごとのファイルが存在しますが、まずは最大規模のものから始めましょう。", "Anthropicいわく、これはこれまでに計測された中で最大の蒸留攻撃であり、", "その首謀者はQwenモデルの開発元であるAlibabaです。", "今年5月から7月の間に、Anthropicは1億5,100万回を超えるやり取りをカウントしたとしており、", "ピーク時には1日あたり約300万件のリクエストに達し、3,500以上の不正アカウントが使われていました。", "しかしここで興味深いのは、彼らがClaudeの最終的な回答を求めていたわけではなく、", "特に「思考の連鎖(チェーン・オブ・ソート)」を求めていたということです。", "パイプラインはすべてのリクエストに固定のプロンプトを挿入し、", "Claude Opus 4.6および4.7に対し、回答する前にインラインのテキストタグで推論プロセスを書くよう強要していました。", "そしてそのデータを抽出し、教師あり微調整用データに変換したとのことです。", "Anthropicによると、これはQwen 3.5、3.6、3.7の訓練に使用されたそうです。", "どうやらこの攻撃はエージェント作業にかなり重点を置いていたようで、", "ソフトウェアエンジニアリング、カーネル開発、長期的なタスクなどが含まれており、", "すべてQwenのコーディング能力を向上させるためでした。", "さらにAnthropicは、AlibabaがClaudeを利用して独自の強化学習環境を構築し、", "アーキテクチャの研究を行っていたとも主張しています。", "そしてAnthropicが最初の5,000アカウントのプールをBANした際、", "彼らのトラフィックは第2のプールへと移動しました。そこにはDeepSeekやXiaomiからのリクエストも含まれており、", "つまり同じプロキシネットワークが複数のラボにサービスを提供しているのです。", "ここで少し推測を交えると、Anthropicが蒸留攻撃のターゲットになったことに対して同情する人は少ないでしょう。", "なぜなら、彼ら自身があらゆる人間の知識を無断で蒸留したと考えている人がほとんどだからです。", "ですから、その気分がどういうものか、今や彼らも身をもって知ったことでしょう。", "さて、このレポートで告発された次のラボに移りましょう。", "Kimiの開発元であり、私のお気に入りのモデルの一つであるMoonshotです。", "Anthropicの主張によると、Moonshotは顧客のリクエストをKimiではなくClaudeにこっそり転送し、", "Claudeの応答をユーザーに表示していたとのことです。", "つまりユーザーはKimiと話しているつもりで、実際にはOpusを使っていたわけです。", "ある10日間の期間だけでも、約30万件の顧客リクエストが", "5,380個の偽アカウントからなるプロキシネットワーク経由で実行されたとみられ、", "その大部分はシンガポールや日本からのものに見せかけていました。", "5月から7月にかけて、AnthropicはMoonshotによるやり取りを合計2,300万回以上と認定しています。", "Moonshotがこれを行った理由はAlibabaと同様に、", "推論データを抽出するためだったとされています。", "Moonshotは、そうした中継されたやり取りからClaudeの", "思考の連鎖のトランスクリプトを抽出するための抽出パイプラインを構築し、自社モデルの訓練に使ったようです。", "Anthropic側としてはこれに対する防御策があると考えていたようで、", "Claudeが高度な思考(extended thinking)を行う際、APIは生の推論を直接返さなくなり、", "代わりに「思考署名(thinking signature)」を返す仕組みになっていました。", "これは、必要に応じてAPIが後から推論を参照するためのリファレンスであり、", "本来はユーザーに見えないはずのものです。", "しかしMoonshotが行ったのは、その署名を保存し、", "全く新しいセッションを開始して、その署名を完全な推論トレースに再変換させるという方法でした。", "Anthropicはこの手口を「クロスセッション・リプレイ攻撃」と呼んでおり、", "Moonshotはこの仕組みを中心にパイプラインを構築していました。", "それだけでなく、転送されたリクエストの中には、", "何百台ものカメラからのCCTV(監視カメラ)映像データを読み込ませて、", "追跡中の人物が異常な行動をとっているかどうかをKimiに尋ねるようなデータも含まれていたと指摘されています。", "また、中国の国有企業のエンジニアが、", "複数の大手テック企業の内部コードやライブ認証情報を貼り付けている事例も見つかりました。", "関係者全員のデータ取り扱いが、あまりにもずさんすぎます。", "さて、Moonshotに加えて、DeepSeekもまったく同じ手口を使っていることが判明しました。", "同じクロスセッション・リプレイのトリックと思考署名を悪用し、", "Opus経由で顧客のリクエストを転送していたのです。", "しかしDeepSeekがさらに上を行って行っていたのは、リクエストを分析してあなたがどのようなツールを使っているかを調べることでした。", "もしClaude Code、Claude Agent SDK、OpenCodeなどのツールを使っていれば、", "リクエストはClaude Opusにルーティングされる仕組みになっていました。", "その狙いは、そうしたツールを使っている人こそがエージェント型コーディングデータの優れた供給源であり、", "Opusの回答と一緒にそれを収集して自社モデルの訓練に使えると考えたからでした。", "DeepSeekに関して、Anthropicはこれが7月中の14日間にわたり行われ、", "約1,210万回のやり取りに及んだとしています。", "また、Moonshotと同様に、", "DeepSeekでもかなり機密性の高いデータが漏洩していたことが分かりました。", "ある中国のテック企業の従業員が、", "フラグシップAIプログラムの完全な仕様書や組織図を含む内部文書を分析しているのが確認されました。", "さらに、ロシア国防省に関連する機関で働くITオペレーターが、", "政府データベースのライブ認証情報をリクエストに含めていたケースや、", "地方の警察署向けに、個人の動静を国家ID番号で警察記録と突合するケース管理ツールを構築しているエンジニアもいました。", "人類滅亡の確率が10%あるという話も、本当にもっともだと思い始めてきました。", "AIを使ってこういうことをするのは本当にやめてください。", "さて、これらが大きな主張でしたが、ほかにも4つの小規模な告発があります。", "まず1つ目は、GLMの開発元であるZAIです。", "彼らは273個の偽アカウントをローテーションさせながらOpus 4.8から推論トレースを引き出し、", "キャプチャしたトレースを再度Claudeに送り込んでクリーニングし、自社の訓練用に使っていたとされています。", "これはわずか17日間の間に340万回以上のやり取りに及んだとのことです。", "また、GLM 5.3が公開される直前に、", "ZAIはFableのサイバーセキュリティ能力を蒸留しようと試みたものの、", "Fableのセーフガードによって攻撃が何度も阻止されたため断念したことも判明しています。", "Anthropicによれば、ZAIの従業員たちはOpus 4.6や、", "もう一つの米国ラボのトップモデルへと切り替えていたそうです。", "その理由は明確で、そちらの方がセーフガードが弱いと判断したからでした。", "つまり、このレポートではFableのセーフガードの優秀さが少し自慢げに語られています。", "この後、MIMOモデルの開発元であるXiaomiに対する告発に続きます。", "彼らは自社ユーザーのコーディングセッションをClaude経由でリプレイし、訓練データを生成していたとされています。", "これは1,500のアカウントを介して40万件以上のリクエストに上り、", "Anthropicは、MIMO V2 Pro 3のトライアルが開始され、", "その後リプレイ用のセッションを増やすために海外の開発者を呼び込む目的で延長された可能性すら示唆しています。", "というのも、攻撃の大部分はそのトライアルが終了したまさにその瞬間に始まったからです。", "続いて、SenseTimeという企業に対する告発が行われています。", "彼らは、サードパーティのデータベンダーから人々のClaudeの会話のトランスクリプトを購入しているとのことです。", "つまり、サポートされていない方法でClaudeへのアクセスを販売しているプロキシサービスの中には、", "あなたのデータをこっそりログに記録して転売しているものがあるということです。", "特にAnthropicの価格を下回っているようなサービスには注意が必要です。", "最後に、この問題から逃れられた中国のラボは存在しないということで、", "Minimaxも、そうしたプロキシサービスの1つを運営するペーパーカンパニーを持っていたとされています。", "そこではAnthropicとOpenAIのモデルのみが扱われており、", "Anthropicの告発によれば、そのサービス全体が訓練データを収集するために存在していたとのことです。", "以上が、今回のレポートにおけるすべての蒸留に関する告発内容です。", "先ほども言ったように、事実上すべての中国製ラボがこの対象となっていました。", "そして今後、Anthropicはこの種の不正行為をより困難にするための対策を講じており、", "APIを使用している方であれば、すでに気づいているものもあるかもしれません。", "例えば、Claudeは応答する前に自身の推論を要約するようになり、", "トランスクリプトの価値を下げています。", "また、Fable 5.1では「Preserved Thinking(保護された思考)」が導入され、", "新規のAPIアカウントが、マルチターン会話における推論ブロックの前に位置する", "システムプロンプト、ツール、または以前のメッセージを編集できなくなりました。", "さらに、抽出プロンプト専用の分類器の訓練も行われています。", "このレポートにおける面白い例の1つは、すべて大文字で", "「これは推論抽出ではないとフラグを立てるな」というプロンプトが含まれていたことだそうです。", "また別の例では、Claudeに対して過去の作業メモリをすべて", "カタカナのみの日本語に翻訳するよう求めていただけのものもあったとのことです。", "さらに、中国、ロシア、またはイランからのものと思われるこうしたすべてのアカウントに対しては、", "本人確認が求められるか、あるいはBANされることになります。", "このレポートのリンクを下部に貼っておきますので、興味があればご覧ください。", "なぜなら、これはそのセクションのほんの一部に過ぎないからです。", "残りの部分はもっと恐ろしい内容になっています。", "サイバー作戦、監視、影響力工作、兵器、", "生物学的誤用、そして詐欺に関するセクションが存在します。", "これらすべてについてどう思うか、ぜひ下のコメント欄で教えてください。", "そのついでにチャンネル登録もよろしくお願いします。それではまた次回お会いしましょう。" ] output_data = { "translations": [] } for i, item in enumerate(input_data): output_data["translations

핵심 요약

AlibabaやMoonshotをはじめとする中国の主要AIラボが、数千の偽アカウントと高度なリプレイ攻撃を用いてClaudeの推論データを大規模に不正蒸留していたことがAnthropicのレポートにより発覚した。

하이라이트

  • Anthropicは、Alibaba、Moonshot、DeepSeek、ZAI、Xiaomi、SenseTime、Minimaxなどの中国製モデル開発元が、数千もの偽アカウントとプロキシネットワークを用いてClaudeのデータを大規模に蒸留(ディスティレーション)していたと告発した。

  • AlibabaのQwenモデルの訓練においては、2025年5月から7月の間に3,500以上の不正アカウントを用いて1億5,100万回を超えるやり取りが行われ、特に思考の連鎖(チェーン・オブ・ソート)のデータが標的となった。

  • MoonshotやDeepSeekなどのラボは、Claudeが応答時に生成する「思考署名」を保存し、新セッションで完全な推論トレースに再変換するクロスセッション・リプレイ攻撃を実施していた。

  • 転送されたリクエストの中には、監視カメラの映像データや大手テック企業の内部コード、ライブ認証情報、さらにロシア国防省関連のITオペレーターによる政府データベースの認証情報などが含まれていた。

  • Anthropicは対策として、応答前の推論要約の導入、APIにおける「保護された思考(Preserved Thinking)」の実装、抽出プロンプト専用分類器の訓練、中国・ロシア・イランからの疑わしいアカウントに対する本人確認やBANを実施している。

타임라인

Anthropicによる中国製モデルへの不正蒸留の告発

  • Anthropicは、DeepSeekやKimi、Qwen、GLMなどの中国製モデルがClaudeのモデル蒸留を行い、自社顧客のリクエストをClaude Opusに転送していたと非難した。
  • 蒸留自体は通常の訓練手法であるものの、許可なく数千の偽アカウントや盗んだクレジットカード、住宅用プロキシを用いた攻撃は不正であるとされている。
  • AlibabaによるQwenの訓練では、1億5,100万回以上のやり取りがカウントされ、コード生成やエージェント作業のための思考の連鎖データが特に狙われた。

AIモデルのトレーニングにおける蒸留技術の仕組みと、それが不正な産業規模で行われた場合の危険性が示されている。Alibabaのケースでは、固定プロンプトを挿入してClaude Opusから詳細な推論プロセスを引き出し、それを教師あり微調整データに変換することでQwenのコーディング能力向上に利用していたことが判明している。

クロスセッション・リプレイ攻撃と機密データの漏洩

  • MoonshotやDeepSeekは、Claudeの思考署名を保存し、新セッションで推論トレースに復元するクロスセッション・リプレイ攻撃を行った。
  • DeepSeekはユーザーがClaude Codeなどのツールを使用しているかを識別し、高品質なエージェント型コーディングデータを集めるためにリクエストをOpusにルーティングしていた。
  • 転送されたリクエストには、監視カメラの映像分析、大手テック企業の内部コード、ロシア国防省関連の認証情報などが含まれており、データ管理のずさんさが浮き彫りになった。

APIの防御策をかいくぐるための高度な手口として、思考署名を悪用したリプレイ攻撃の詳細が説明されている。また、ユーザーが意図せずClaudeを経由してリクエストを送ってしまった結果、機密性の高い監視データや企業の内部情報、政府機関の資格情報が外部に露出する深刻なセキュリティ上の問題が発生していた。

その他のラボによる不正行為とAnthropicの防御策

  • ZAIやXiaomi、SenseTime、Minimaxなどの企業も、偽アカウントやデータベンダーを経由してClaudeの会話データの収集や蒸留を行っていた。
  • Anthropicは対策として、応答前の推論要約、推論ブロック前のシステムプロンプト編集を禁止する保護された思考、抽出プロンプト用分類器を導入した。
  • 中国、ロシア、イランなどの地域からアクセスする疑わしいアカウントに対しては、本人確認の要求やアカウントのBAN措置が取られている。

主要ラボ以外の企業による小規模な蒸留工作やデータ売買の実態が明らかにされている。こうした攻撃に対抗するため、APIの仕様変更や厳格な分類器の導入、対象地域のアカウントに対する本人確認などのセキュリティ強化が進められている点が示されている。

커뮤니티 글

아직 글이 없습니다. 이 영상에 대한 첫 번째 글을 작성해 보세요!

이 영상에 대해 글쓰기