为 DeepSeek V4.0 等单模态大模型装上眼睛和耳朵 —— 本地视觉+音频 MCP 服务器,基于 Ollama + MiniCPM-V 4.6 + faster-whisper,图片描述·视频分析·语音转文字
-
Updated
Jun 26, 2026 - JavaScript
为 DeepSeek V4.0 等单模态大模型装上眼睛和耳朵 —— 本地视觉+音频 MCP 服务器,基于 Ollama + MiniCPM-V 4.6 + faster-whisper,图片描述·视频分析·语音转文字
Codex plugin that gives text-only models (e.g. DeepSeek) image understanding via Qwen Vision API · 为非多模态模型提供看图能力的 Codex 插件
Add a description, image, and links to the image-description topic page so that developers can more easily learn about it.
To associate your repository with the image-description topic, visit your repo's landing page and select "manage topics."