Local multimodal vision & speech for text-only LLM agents (Codex, Claude Code, Cursor, Cline, Gemini CLI, and similar). OCR, image/video understanding, and audio transcription run entirely on your machine via Ollama + FunASR, with a CLI and MCP server.
Or install with pip: pip install local-agent-senses
README Excerpt
Local vision, video and speech tools for text-only LLM agents. The default backend is Ollama on your own machine; media is not uploaded unless you explicitly configure an OpenAI-compatible endpoint. 中文说明见 [README.zh-CN.md](README.zh-CN.md)。 中文版本请见 [README.zh-CN.md](README.zh-CN.md)。 - Image understanding and verbatim OCR/transcription.