README Excerpt
Lets a **text-only LLM read images and PDFs** by delegating to a vision model running on your own GPU, through `llama.cpp`. Built for the case where your coding model has no vision at all — GLM on the Z.AI coding plan, DeepSeek, most local models. The coding model stays where it is; this server becomes its eyes.
Tools (20)
VISION_ALLOWED_ROOTSVISION_API_KEYVISION_CONTEXTVISION_EXTRA_ARGSVISION_GPU_LAYERSVISION_HTTP_HOSTVISION_HTTP_PORTVISION_HTTP_TOKENVISION_IDLE_MSVISION_MANAGEDVISION_MAX_EDGEVISION_MAX_FILE_BYTESVISION_MAX_PAGESVISION_MAX_TOKENSVISION_MMPROJ_PATHVISION_MODEL_PATHVISION_NODEVISION_PDF_DPIVISION_SERVER_BINVISION_SERVER_URL