DeepSeek Harness Adds Multimodal Image Workflows

๐กBuild agents that retain and route images across goals, plans, sessions, MCP/ACP, and nested tool calls.
โก 30-Second TL;DR
What Changed
Adds the DeepSeek-V4-Flash-Vision-Exp multimodal adapter
Why It Matters
This release makes DeepSeek Harness more suitable for agents that need to combine visual context with planning, goals, files, and persistent sessions. Developers can build richer multimodal workflows without handling every image attachment manually at the application layer.
What To Do Next
Install DeepSeek Harness v0.1.1-rc.1 and prototype an image-aware /plan workflow using the DeepSeek-V4-Flash-Vision-Exp adapter.
Key Points
- โขAdds the DeepSeek-V4-Flash-Vision-Exp multimodal adapter
- โขSupports native image requests and image inputs for /goal and /plan
- โขThe @ menu can reference files and sessions
- โขMCP/ACP supports persistent image attachments
- โขPTC Mode can forward nested images
๐ง Deep Insight
Background and context from public sources โ not the original article. 14 sources cited.
๐ Enhanced Key Takeaways
- โขDeepSeek Harness is built on the 'Cordis' meta-framework, which utilizes a modular, plugin-first architecture for swappable tools and agent loops.
- โขThe new vision model tokenizes images at a fixed rate of 384 tokens per image, maintaining parity with standard V4-Flash pricing.
- โขA new Files API allows for persistent image storage via file_id, reducing bandwidth and costs by eliminating redundant uploads in multi-turn sessions.
- โขThe v0.1.1 update transitioned previously bundled components like Claude Code and Codex into optional 'Profile Bundles' to reduce core installation bloat.
- โขDeepSeek Harness achieved rapid market traction, reaching over 160,000 GitHub stars within the first week of its August 13, 2026, open-source release.
๐ Competitor Analysisโธ Show
| Feature | DeepSeek Harness | Claude Code | Codex |
|---|---|---|---|
| Architecture | Modular (Cordis) | Bundled | Bundled |
| Pricing | 30-100x cheaper (est) | Standard API | Standard API |
| Multimodal | Native (V4-Flash) | Yes | Yes |
| Customization | High (Plugin-first) | Low | Low |
๐ ๏ธ Technical Deep Dive
- Built on the Cordis meta-framework for component modularity.
- Implements a tokenization strategy of 384 tokens per image for the V4-Flash-Vision-Exp model.
- Uses a Files API for persistent storage and reference by file_id to optimize multi-turn context windows.
- Supports PTC (Persistent Task Context) Mode for handling nested image forwarding in autonomous loops.
- Architecture supports on-demand installation of external agent tools via Profile Bundles.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (14)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.