來源Reddit r/LocalLLaMA•較早收集於 3h
23,759 跨模態提示注入酬載開源

#prompt-injection#multimodal-security#red-teamingbordair-multimodal-v1distilbertowasp-llm-top-10
💡分割酬載規避多模態防禦—LLM 安全必看(48字元)
⚡ 30 秒速覽
有什麼變化
23,759 個酬載跨文字+影像+文件+音訊模態分割
為什麼重要
凸顯多模態 LLM 漏洞,敦促統一跨通道偵測。對建構強健防禦的安全研究者至關重要。
下一步行動
從 GitHub 下載酬載,並對您的多模態 LLM 偵測管線測試。
誰應關注:Researchers & Academics
關鍵要點
- •23,759 個酬載跨文字+影像+文件+音訊模態分割
- •規避 DistilBERT 分類器(每個片段 0.43-0.53 信心)
- •類別:資料外洩、越獄、編碼混淆、多語言
- •組合:文字+影像 EXIF、PDF 中繼資料、超音波音訊、隱藏 PPTX 層
- •僅 JSON 儲存庫,供紅隊評估多模態 LLM 偵測
🧠 深度解析
本篇為 AI 生成分析,非原文內容。
🔑 增強重點摘要
- •The dataset utilizes a 'fragmented-payload' strategy where individual components are designed to trigger low-confidence alerts in standard safety classifiers, effectively bypassing threshold-based filtering systems.
- •The repository includes specific implementations for steganographic embedding, such as hiding malicious instructions within the least significant bits (LSB) of image files and manipulating PDF cross-reference tables to bypass document scanners.
- •Security researchers have identified that the effectiveness of these payloads relies on the 'reconstruction' capability of multimodal LLMs, which aggregate seemingly benign fragments from different modalities into a coherent, malicious prompt during the inference process.
🛠️ 技術深入
- •Payloads are structured in a JSON schema that maps specific modality-based triggers to target LLM architectures, including support for Vision-Language Models (VLMs) and Audio-Language Models.
- •The dataset employs adversarial noise injection techniques specifically tuned to evade DistilBERT-based text classifiers and ResNet-based image safety filters.
- •The audio payloads utilize ultrasonic frequency modulation (above 18kHz) to remain imperceptible to human listeners while remaining detectable by high-fidelity microphone inputs used in voice-enabled LLM interfaces.
- •The document-based payloads leverage hidden metadata fields and non-rendered text layers in PPTX and PDF formats to bypass standard OCR and text-extraction safety pipelines.
🔮 前景展望基於引用來源的 AI 分析
Multimodal safety filters will shift from per-channel analysis to holistic cross-modal fusion architectures.
The success of fragmented payloads proves that independent modality checks are insufficient to detect coordinated, multi-vector attacks.
Standardized red-teaming benchmarks for LLMs will mandate cross-modal injection testing by 2027.
The release of this large-scale dataset establishes a new baseline for evaluating the robustness of multimodal safety defenses against sophisticated obfuscation.
⏳ 時間線
2025-11
Initial research paper published on cross-modal prompt injection vulnerabilities in multimodal LLMs.
2026-02
Development of the automated payload generation framework begins, focusing on fragmenting malicious prompts.
2026-04
Public release of the 23,759-payload dataset on GitHub and announcement on r/LocalLLaMA.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/LocalLLaMA ↗
每週電子報
每週一封,可隨時退訂。