How much faster was Ollama on Linux Mint than on Windows 10?
37.2% faster at generating text: 12.9 tokens per second on Linux Mint 22.3 against 9.4 on Windows 10, a gain of 3.5. Both ran Ollama 0.11.4 with the same TinyLlama Q4_0 file, on the processor alone.
| Measurement | Windows 10 | Linux Mint 22.3 |
|---|---|---|
| Tested on | September 29, 2026 | October 10, 2026 |
| Generation speed (tokens per second, median of 3) | 9.4 | 12.9 |
| First answer from cold (seconds) | 229.8 | 19.8 |
| Model memory (GB, Ollama’s report) | 0.7 | 0.7 |
The laptop is an HP Notebook with a two-core Intel Core i3-5005U, 6 GB of RAM and a hard drive. I benchmarked it on Windows 10 Home 22H2 on September 29, then replaced Windows with Linux Mint 22.3 using our install guide and ran the same test on October 10. The model file’s digest matched on both systems, so the weights were identical.
Why did Ollama take so long to start on Windows 10?
Mostly startup. Ollama’s runner took about 167 seconds to get ready on Windows 10, inside a first answer that took 229.8 seconds in total. On Mint the whole first answer took 19.8 seconds.
The cold time counts everything from an unloaded model to a finished answer: starting the runner, reading the model from disk and writing the reply. So the 91.4% drop isn’t a measure of generation speed, and I can’t say how much of it the operating system caused. On Windows, other apps and Ollama’s tray app were running, and the test doesn’t reboot or clear the disk cache between runs.
What model should I run on a laptop without a GPU?
Start with a small one. TinyLlama (1.1B parameters, Q4_0) needed 0.7 GB once loaded and ran at 12.9 tokens per second on this laptop with 6 GB of RAM.
For scale, my 2012 Xeon E5-2665 desktop runs the same model at 27.1 tokens per second on its CPU with the same Ollama version, so the laptop on Mint reached 47.6% of a desktop with eight cores. Larger models are slower on any CPU: the CPU-only benchmark has measured speeds from 1B to 33B on that desktop.
Should I switch an old laptop to Linux to run local AI?
If it runs Windows 10, it’s worth trying. On this laptop, Linux Mint 22.3 gave the same model 3.5 more tokens per second, and Mint runs from a USB drive before you install anything.
Back up first: replacing Windows erases the disk. Our verified Linux Mint download and six-step install guide cover the USB, the install and the Wi-Fi driver fix this HP needed.
How do I test my own laptop?
Install Ollama, pull TinyLlama and run bench.py from local-ai-benchmarks. Run it before and after a change, with the same model file, and compare the warm speed.
$ curl -O https://raw.githubusercontent.com/Arynwood-Technology/local-ai-benchmarks/main/bench.py $ ollama pull tinyllama $ python3 bench.py tinyllama --modes cpu --runs 3 --num-predict 200 --csv my-laptop.csv
Close other apps, keep the charger connected and let the desktop sit idle for two minutes first. Check that ollama list shows the same digest on both systems; the tag alone doesn’t guarantee the same file. If you test a different laptop, send in your numbers.
How I tested
HP Notebook: Intel Core i3-5005U (2 cores, 4 threads, 2.00 GHz), 6 GB RAM, Intel HD Graphics 5500 (not used: num_gpu: 0), SATA hard drive. Windows 10 Home 22H2 (build 19045), Balanced power plan, September 29, 2026. Linux Mint 22.3 with Cinnamon 6.6.4 and kernel 6.14.0-37-generic, October 10, 2026. Ollama 0.11.4 and TinyLlama Q4_0 on both, with the same model digest (ID 2644915ede35 in ollama list). bench.py, unmodified: one cold request after unloading the model, then three warm requests at temperature 0 and seed 42, with a 200-token cap. It’s a paired test of one laptop on two dates, not a lab isolation of the operating system, and three warm runs from one session can’t show run-to-run variation. Full paper: Linux Mint ran a local AI model 37% faster than Windows 10. Report, raw data (CC BY 4.0) and script (MIT): DOI 10.5281/zenodo.23290682. Method: How we test.
- October 10, 2026: first published.