← All writing
Strix Halo

Ollama Vision Is Slow on Your AMD APU? It’s Running the Image Encoder on CPU — Here’s the 4x Fix

Originally published on Medium. This archived article reflects the projects, opinions, and versions at the time of publication. View the Medium original ↗

In this article
The symptomHow to diagnose itWhy Ollama does thisThe fix: run llama-server directly with the projector on the GPUGotcha: if your vision model is a “thinking” model, you’ll get empty answersMake it a persistent “vision sidecar”The benchmarkTwo bonus tips that stackThe honest caveat — test your own hardware
Tested on a Strix Halo (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151). If you’re running qwen3-vl or any vision model on an iGPU/APU and it’s crawling, this is almost certainly why.
Tested on a Strix Halo (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151). If you’re running qwen3-vl or any vision model on an iGPU/APU and it’s crawling, this is almost certainly why.

I run local vision-language models on an AMD Strix Halo mini PC for a camera project — feed it a frame, ask it what’s going on. The text models scream on this thing. But the moment I pointed a vision model (qwen3-vl:4b) at a 4K camera frame, inference fell off a cliff: ~40 seconds per image. Worse, while it churned, one CPU core group sat pegged near 1000% (about 10 cores) and the GPU was basically idle for the image step.

A 128GB unified-memory APU with a perfectly capable iGPU, and it’s grinding images on the CPU. Something was off. Here’s what’s happening and how to get it back.

The symptom

That last point is the tell. If text is fast but images are slow, the language model is on the GPU — it’s the vision encoder that isn’t.

How to diagnose it

Watch the Ollama logs while you run a vision request. The smoking gun is this line:

disabling multimodal projector offload  reason=shared-memory-gpu

Right after that, Ollama launches its internal llama-server with the flag --no-mmproj-offload. Translation: the CLIP/vision projector — the image encoder, the "mmproj" — is being forced onto the CPU, while the language half stays on the GPU. That split is exactly why you see a massive CPU spike for the image and an idle GPU.

You can confirm the CPU side with htop or top during a request: you'll see one process eating ~1000% CPU for the duration of the encode.

Why Ollama does this

It’s not a bug, it’s a deliberate, hardcoded heuristic. Ollama disables GPU projector offload for any shared-memory GPU — that’s every iGPU and APU, where the GPU borrows system RAM instead of having dedicated VRAM. The reasoning is OOM protection: on some shared-memory setups, loading CLIP onto the GPU can blow up memory or produce garbage, so Ollama plays it safe.

The frustrating part: it’s a blanket rule. It doesn’t matter that you have 128GB of unified memory sitting mostly free — if the GPU is shared-memory, the projector goes to CPU. And as of Ollama 0.30.11, there is no environment variable to override it. (See Ollama GitHub issues #10889 and #13742 — this has bitten a lot of APU users.)

The fix: run llama-server directly with the projector on the GPU

Here’s the good news. Ollama is built on llama.cpp, and it bundles the llama-server binary and has already downloaded your model. You can run that same binary yourself, on a different port, with the projector offload enabled — and point your app at it.

First, find the bundled binary and the model files. The binary typically lives here:

# The llama-server Ollama ships with:
ls /usr/local/lib/ollama/llama-server
# Find the model files Ollama already pulled.
ls ~/.ollama/models/manifests/registry.ollama.ai/library/qwen3-vl/
ls ~/.ollama/models/blobs/

The model is stored as content-addressed blobs (files named sha256-...), and the manifest JSON maps the model name to those blobs. Pull the model-layer digest from the manifest:

python3 -c "import json; m=json.load(open('$HOME/.ollama/models/manifests/registry.ollama.ai/library/qwen3-vl/4b')); print(next(l['digest'].replace(':','-') for l in m['layers'] if l['mediaType']=='application/vnd.ollama.image.model'))"
# -> sha256-...   (the file lives in ~/.ollama/models/blobs/)

Important — qwen3-vl ships as one combined GGUF. The vision projector is packed into the same file as the language weights, so you pass the same blob to both --model and --mmproj — there is no separate mmproj file to hunt for. (Some other vision models do split the projector into its own layer; if yours does, your manifest will show a second projector/mmproj layer and you'd point --mmproj at that one. For qwen3-vl:4b, it's one file.)

Then launch your own vision server with the projector on the GPU. I reused the llama-server binary and ROCm backend Ollama already bundles, so the three env vars below point at them — if you built llama.cpp yourself, drop the env lines and use your own binary:

BLOB=~/.ollama/models/blobs/sha256-<model-layer-digest>
GGML_BACKEND_PATH=/usr/local/lib/ollama/rocm_v7_2/libggml-hip.so \
LD_LIBRARY_PATH=/usr/local/lib/ollama:/usr/local/lib/ollama/rocm_v7_2 \
ROCR_VISIBLE_DEVICES=0 \
/usr/local/lib/ollama/llama-server \
  --model "$BLOB" \
  --mmproj "$BLOB" \
  --mmproj-offload \
  --host 127.0.0.1 --port 8081 \
  -c 8192 -np 1 \
  --flash-attn on \
  --image-min-tokens 1024 \
  -b 1024 -ub 1024 --context-shift --keep 4

The flag that does the work is --mmproj-offload — the one Ollama refuses to set for you. Note --model and --mmproj point at the same blob (combined GGUF, see above). If you find the language layers aren't all landing on the GPU, add --n-gpu-layers 999. Then point your app's OpenAI-compatible calls at http://127.0.0.1:8081.

Gotcha: if your vision model is a “thinking” model, you’ll get empty answers

This one cost me an hour, so let me save you the trouble. qwen3-vl:4b is a thinking variant — left to the default template it opens every reply with a long <think>…</think> reasoning monologue before the actual answer. That's harmless until you add a token cap to keep responses short (which you'll want to). Then the cap gets eaten by the hidden thinking and you get a truncated, empty answer. (Ollama's think: false doesn't save you — it just hides the monologue in a separate field; the tokens are still spent.)

The fix is a tiny chat template that prefills a closed, empty think block on the assistant turn, so the model skips the monologue and answers directly. Save this as nothink.jinja:

{%- for message in messages -%}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>\n' -}}
{%- endfor -%}
{%- if add_generation_prompt -%}
{{- '<|im_start|>assistant\n<think>\n\n</think>\n\n' -}}
{%- endif -%}

Then add --jinja --chat-template-file /path/to/nothink.jinja to the launch command above. Clean, direct, single-digit-second descriptions, no output-stripping needed. If your vision model isn't a thinking model, skip this entirely.

Make it a persistent “vision sidecar”

You don’t want to babysit a terminal. Wrap that command in a small systemd service so it stays running and the model stays warm in memory. Run it on its own port and leave Ollama alone for everything else — text generation, embeddings, model management. Ollama handles those untouched; your sidecar just owns vision. Best of both worlds, and you skip the cold model-load on every call.

A minimal unit is just ExecStart= pointing at the command above, plus Restart=always. Enable it and forget it.

The benchmark

I ran this apples-to-apples: two isolated llama-server instances, same 4K frame, identical except for the offload flag.

Run                                     Projector   Cold (first)   Warm (cached)
--no-mmproj-offload  (Ollama default)   CPU         39.8 s         1.0 s
--mmproj-offload     (the fix)          GPU         10.0 s         1.0 s

About 4x faster on a cold image encode. The output was coherent and identical in quality with GPU offload — so on this gfx1151 box, GPU CLIP is perfectly stable. (Warm/cached calls are ~1s either way, because at that point you’re not re-encoding.)

Two bonus tips that stack

1. Downscale the image before you send it. Image-encode time scales hard with resolution, and a 4K frame is overkill for most vision tasks. Resize to roughly 720p / ≤1280px on the long edge before the request. You’ll often shave a big chunk off the encode with no meaningful loss in answer quality — the model isn’t reading license plates, it’s telling you there’s a person on the porch.

2. Keep the model resident. That’s the sidecar point again: a persistent server means the weights stay loaded, so you don’t pay a cold model-load on top of the encode for every single call. These two stack with the GPU-offload fix.

The honest caveat — test your own hardware

Ollama disables GPU projector offload on shared-memory GPUs for a reason. Some APUs really may OOM or spit out garbage with CLIP on the GPU. I’m not telling you Ollama is wrong to be cautious — I’m telling you it’s too cautious for this box.

So:

Tested on Strix Halo gfx1151 — your mileage may vary. But if you’ve got a capable iGPU sitting idle while one CPU sweats through every image, this is very likely your 4x.

More from the studio

Explore all 13 articles →See the current work →