Search

Tag: #multi-modal27 results

Meissa: Lightweight Offline Medical AI Agent

Meissa: Lightweight Offline Medical AI Agent

Meissa is a 4B-parameter multi-modal medical LLM that enables offline agentic capabilities by distilling trajectories from frontier models like GPT. It features unified trajectory modeling, three-tier stratified supervision for strategy selection, and prospective-retrospective learning for execution. Trained on 40K trajectories, it matches or exceeds proprietary agents across 13 medical benchmarks with 25x fewer parameters and 22x lower latency; fully open-sourced.

ArXiv AIResearchMar 11#medical-ai#agentic-llm#multi-modal
Perceptron Mk1: 80-90% Cheaper Video AI

Perceptron Mk1: 80-90% Cheaper Video AI

Perceptron launched Mk1, a flagship video analysis reasoning model 80-90% cheaper than Anthropic's Claude Sonnet 4.5, OpenAI's GPT-5, and Google's Gemini 3.1 Pro. Priced at $0.15 per million input tokens and $1.50 output, it excels in spatial and video benchmarks like EmbSpatialBench (85.1) and VSI-Bench (88.5). A public demo is available for testing.

Lemonade v10 Launches Linux NPU Support

Lemonade v10 Launches Linux NPU Support

Lemonade v10 introduces Linux NPU support alongside robust multi-modal capabilities including image generation/editing, transcription, and speech synthesis. It supports Ubuntu, Arch, Debian, Fedora, and Snap with a unified base URL for apps. A control center app aids model and backend management, with community challenges offering AMD laptops.

Reddit r/LocalLLaMACommunityMar 13#npu#multi-modal#local-ai
Qwen Code SDK TypeScript v0.1.8 Released

Qwen Code SDK TypeScript v0.1.8 Released

The Qwen Code SDK v0.1.8 introduces multi-modal input support, improved error handling for API providers, and enhanced subagent tracking. This release also includes significant stability fixes for CLI execution and IDE companion integrations.

Qwen (GitHub Releases: qwen-code)MediaJul 14#sdk#multi-modal#cli
Page 1 of 3