RTX 5070 Ti 16GB owners: use Q4_K_XL and offload everything (-ngl 999) — the model is ~15GB and fits.
Keep context ≤ 16K or offload KV cache to RAM if you need long context.
Expect 25-50 tok/s depending on overclock and power limits.
Route 3 — No GPU: unified-memory machines
Ryzen AI Max+ 395 mini PCs (128GB): 9-16 tok/s generation, 70-135 tok/s prompt processing. Community reports 45-73% MTP draft acceptance, which keeps generation usable.
Mac Mini M4 Pro 64GB: 15-25 tok/s, nearly silent, ~25W average.
Bottom line
Budget
Pick
Expected tok/s
~$500-600
Used RTX 3090
40-80
Have a 5070 Ti already
Q4_K_XL + full offload
25-50
~$800-1000, want silence
Mac Mini M4 Pro 64GB
15-25
~$900, no GPU desk PC
Ryzen AI Max+ 395 128GB
9-16
If you only run short prompts and don't mind slower generation, the unified-memory route is the cheapest total cost of ownership — no GPU, no power draw, no resale gamble.