來源虎嗅•較早收集於 6m
理解 AI 開發中的 Harness Engineering
#ai-reliability#workflow-automation#system-designharness-engineeringopenaianthropichashicorplangchain
💡了解為何頂尖 AI 專家正將重心從提示詞工程轉向「Harness Engineering」,以顯著提升 AI 的可靠性。
⚡ 30 秒速覽
有什麼變化
AI 性能不僅取決於模型本身,還取決於「Harness」(控制系統與工具)。
為什麼重要
這一轉變強調,構建生產級 AI 不再僅僅是提示詞工程,而是更多地關於設計具備彈性、自動化的基礎設施,以約束和引導模型的行為。
下一步行動
審查您當前的 AI 工作流程,將重複的手動提示詞修復替換為永久的系統級配置,例如自定義指令或自動化驗證腳本。
誰應關注:Developers & AI Engineers
關鍵要點
- •AI 性能不僅取決於模型本身,還取決於「Harness」(控制系統與工具)。
- •核心目標是通過修改運行環境來永久解決重複問題,而不僅僅是重新提示。
- •根據 Harness 設計的質量,AI 性能差異可達 6 倍之多。
- •自定義指令、知識庫和自動化檢查迴路等常見做法,本質上都是 Harness Engineering。
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 20 個來源。
🔑 增強重點摘要
- •Harness engineering is a distinct architectural layer from prompt engineering and context engineering, specifically designed for multi-session, multi-agent problems by introducing structured environments, context resets, and phase gates.
- •The discipline operates through reinforcing layers, including preventive controls like constraint harnesses (e.g., lint rules, architectural configurations) and corrective controls such as feedback loops and quality gates, to ensure reliable AI agent performance.
- •Pioneering implementations include OpenAI's internal experiment, where a software product was built with zero manually-written code using Codex, and Spotify's Honk system, both leveraging harness engineering for agent reliability and scalable code generation.
- •The term "harness engineering" was coined by Mitchell Hashimoto in early February 2026, quickly gaining recognition with subsequent engineering articles from OpenAI and Anthropic expanding on the concept.
🛠️ 技術深入
- Layered Control Systems: Harness engineering employs a multi-layered approach, starting with preventive controls (feedforward) like constraint harnesses that reduce an agent's solution space before generation. These can include rules files, architectural lint configurations, and type systems.
- Corrective Mechanisms: Beyond prevention, corrective controls such as feedback loops enable self-correction without human intervention, and quality gates enforce standards that initial layers might miss.
- Standardized Constraints: OpenAI's production system, for instance, enforces "taste invariants"—a set of mechanical rules encoded directly into the repository and enforced as hard CI failures, not just warnings.
- Cross-Tool Specifications: The AGENTS.md spec, released in August 2025, serves as an open standard for cross-tool conventions in the AI coding ecosystem, facilitating monorepo-scale constraint composition.
- Advanced Control Algorithms: Research, such as that from Johns Hopkins University, combines neural networks (specifically an Actor-Critic framework) with mathematical tools like Zubov's Partial Differential Equation (PDE) to create more reliable control systems for complex machines, expanding their operational stability (Region of Attraction).
- Continuous Feedback Loops: AI feedback loops are dynamic processes where an AI system receives performance feedback, adjusts its algorithms (e.g., via backpropagation in neural networks), and then receives further feedback, enabling continuous learning and adaptation.
- AI Governance Frameworks: These frameworks provide systematic approaches with policies, procedures, and guidelines for responsible AI development, deployment, and operation, addressing ethical, legal, and operational challenges like algorithmic bias, data privacy, and model transparency.
- AI Controls: These are comprehensive sets of policies, processes, technical mechanisms, and governance frameworks designed to manage and monitor AI systems across their entire lifecycle, including access controls, data protection, deployment strategies, inference security, and continuous monitoring.
🔮 前景展望基於引用來源的 AI 分析
AI development will increasingly shift from direct code writing to designing and maintaining AI agent environments and control systems.
The rise of harness engineering and agentic platforms suggests a future where engineers primarily build the 'harness' that enables AI agents to reliably generate and manage code and tasks.
Regulatory bodies will integrate 'harness engineering' principles into AI governance frameworks to ensure accountability and trustworthiness.
As AI systems become more autonomous and critical, the mechanisms for control, reliability, and ethical operation inherent in harness engineering will become essential components of compliance and risk management.
The demand for specialized 'Harness Engineers' will emerge as a distinct role in AI teams.
The complexity of designing, implementing, and maintaining robust control systems, feedback loops, and environmental constraints for AI agents will necessitate dedicated expertise beyond traditional prompt or model engineering.
⏳ 時間線
1956
The term "Artificial Intelligence" coined at the Dartmouth Conference, establishing AI as a distinct field of study.
1970s
Emergence of Advanced Process Control (APC) techniques for optimizing complex industrial processes, utilizing sophisticated algorithms and real-time data analysis.
2013-2014
Discovery of "adversarial examples" exposes fundamental vulnerabilities in neural networks, leading to increased focus on AI robustness.
2025-08
The AGENTS.md spec is released as an open standard for cross-tool convention in the AI coding ecosystem.
2025-08
OpenAI begins an internal experiment to build a software product with zero lines of manually-written code, entirely generated by Codex, demonstrating early harness-centric development.
2026-02
Mitchell Hashimoto publishes a blog post coining the term "harness engineering," which is subsequently expanded upon by OpenAI and Anthropic.
📎 來源 (20)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 虎嗅 ↗
每週電子報
每週一封,可隨時退訂。



