qwenv-tools.jinja combines the Qwen2.5 tool-call format with Qwen2.5-VL’s image/video markers. It is selected explicitly by the QwenV service.

Sources:

Keep image markers and tool-call delimiters intact: llama-server uses them to process vision inputs and recognize structured tool calls.