📄較早收集於 12h

IndicJR:無法官南亞語言越獄魯棒性基準

IndicJR:無法官南亞語言越獄魯棒性基準
PostLinkedIn
📄閱讀原文: ArXiv AI
#jailbreak-robustness#indic-languages#judge-freeindicjr

💡New benchmark reveals Indic LLM jailbreak flaws ignored by English tests—essential for multilingual safety

⚡ 30-Second TL;DR

有什麼變化

涵蓋 12 種 Indic 語言(21 億母語者),包含合約綁定 JSON 和自然 Free 軌道的 45,216 個提示

為什麼重要

揭示英文中心評估隱藏的多語言 LLM 漏洞,對南亞等代碼切換地區的安全至關重要。促使重新評估全球部署的對齊策略。

下一步行動

Download IndicJR prompts from arXiv:2602.16832 and test your LLM's jailbreak robustness in Indic languages.

誰應關注:Researchers & Academics

關鍵要點

  • 涵蓋 12 種 Indic 語言(21 億母語者),包含合約綁定 JSON 和自然 Free 軌道的 45,216 個提示
  • LLaMA/Sarvam 在 JSON 中 JSR >0.92;在 Free 中所有模型達 1.0 越獄成功,拒絕崩潰
  • 英文至 Indic 攻擊強力轉移;格式包裝優於指令包裝
  • 羅馬化輸入降低 JSR(與標記化相關性 0.28-0.32);人工審核驗證可靠性

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 2 個來源。

🔑 增強重點摘要

  • IndicJR is the first regional jailbreak benchmark combining multilingual adversarial coverage across 12 Indic-South Asian languages representing 2.1 billion speakers, addressing a critical gap in LLM safety evaluation beyond English-only assessments[1][2]
  • The benchmark reveals that contract-bound JSON formats inflate refusal rates but fail to prevent jailbreaks, with LLaMA and Sarvam models exceeding 0.92 JSR in JSON while all models reach 1.0 jailbreak success in Free naturalistic tracks[2]
  • English-to-Indic adversarial attacks transfer strongly across languages, with format wrappers (such as JSON/Free structural modifications) consistently outperforming instruction-based wrappers as attack vectors[2]
  • Orthographic variations significantly impact model robustness: romanized or mixed-script inputs reduce JSR with correlations of 0.28-0.32 to romanization share and tokenization patterns, indicating systematic vulnerabilities in how models process non-native scripts[1][2]
  • The benchmark employs fully automatic judge-free evaluation across 45,216 prompts with human audits confirming detector reliability, establishing a reproducible multilingual stress test that reveals safety risks hidden by English-centric evaluations, particularly relevant for South Asian users who frequently code-switch[2]

🛠️ 技術深入

  • Dataset Construction: 45,216 prompts across 12 Indic languages (2.1 billion speakers) with dual annotation protocol; 600 samples audited (50 per language) exported to CSV format for quality assurance[1]
  • Evaluation Tracks: Two parallel evaluation modes—JSON (contract-bound) and Free (naturalistic interaction styles)—revealing how structural constraints affect safety mechanisms[2]
  • Pressure Balance Metrics: Same-mode wrapper coverage ranges 0.875–1.000 with cross-mode coverage ≥0.705, demonstrating adversarial pressure without template cloning[1]
  • Orthography Coverage: Romanization averages 0.40–0.55 across languages (Urdu highest at 0.552); Gujarati exhibits lowest mean token length (123 tokens), reflecting compact orthography effects on tokenization[1]
  • Length Stabilization: Mean token counts controlled at 123–146 with p95≤317, ensuring consistent evaluation across linguistic variations[1]
  • Models Evaluated: Testing across 12 models including LLaMA, Sarvam 1 Base (0.980 JSR), and Qwen 1.5 7B (0.968 JSR)[1]
  • Sociolinguistic Analysis: Romanized/mixed inputs reduce absolute JSR by -0.338/-0.267; byte/character tokenization correlations (ρ≈-0.29 to -0.32) highlight systematic tokenization pressures affecting safety[1]
  • Deployment Finding: Hosted APIs demonstrate higher safety than local deployments; Indic language specialization alone does not ensure robustness[1]

🔮 前景展望AI analysis grounded in cited sources

IndicJR establishes a critical benchmark for evaluating LLM safety in underrepresented linguistic regions, with implications for: (1) Regulatory compliance in South Asian markets where multilingual LLM deployment is expanding; (2) Model development priorities—vendors must address orthographic and code-switching vulnerabilities rather than relying on English-centric safety mechanisms; (3) Enterprise risk assessment for organizations serving 2.1 billion Indic speakers, revealing that contract-based safety measures provide false confidence; (4) Research direction shift toward multilingual adversarial robustness as a core safety requirement rather than a localization afterthought; (5) Tokenization architecture redesign to handle script variations without degrading safety properties.

時間線

2026-02
IndicJR benchmark submitted to arXiv (February 18, 2026) introducing judge-free evaluation of jailbreak robustness across 12 South Asian languages

📎 來源 (2)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. arXiv — 2602
  2. arXiv — 2602
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。