Running Ollama Locally with Podman on Fedora Linux

You need dedicate support in the runtime for NPU based inference. Unfortunately people gravitate towards ollama because of its low starting threshold, but it does not include the latest developments in various runtimes, not to mention they have a poor history of crediting other FOSS projects it builds on.

Anyway, to answer your question, your best bet is to try the latest llama.cpp release for your hardware, or if you have AMD hardware you could try lemonade. For both, you would need to do some digging to ensure you have enabled it with the right CLI flag or config option. There’s also FastFlowLM, but I have not tried it yet.