Writing

What I learned running models locally

· 6 min read

I run open models on my workstation: an NVIDIA RTX PRO 4500 Blackwell with 32 GB of VRAM, a Ryzen 9 9950X, and 64 GB of RAM. I wanted to know which jobs could move off hosted models. The answer was fewer than I hoped, and for different reasons than I expected.

The setup

Ollama does most of the serving: gpt-oss 20b, Qwen 3.5 and 3.6 (the 35B mixture-of-experts and the dense 27B), Hermes 3, and qwen2.5vl for images. I also keep llama.cpp built with Vulkan for llama-server. Training runs on TRL, PEFT, bitsandbytes, and Unsloth for QLoRA. faster-whisper and Piper handle speech. LeRobot handles robot-arm policies.

Every experiment below had a scorer before it had a verdict.

Local vs hosted on a real agent

I run a household agent that adds groceries, to-dos, and reminders, and can run a few operator commands. I replayed 20 cases against three models. A deterministic scorer graded tool choice, arguments, mutation safety, and the reply.

Model Where Score Mean latency Critical failures
GPT-5.5 hosted 0.97 5.1 s 1
gpt-oss:20b local 0.87 1.8 s 4
qwen3.6:35b-a3b local 0.74 31.4 s 6

gpt-oss:20b was almost three times faster than the hosted model and 0.10 lower. The misses were in the wrong places. Its critical failures were about writes: retrying an add after a timeout without checking whether the first one landed, mishandling a duplicate, not reporting a failed write honestly, and not preserving an exact command.

The Qwen MoE fit on the card at about 27 GB. Memory was not the problem. Three calls errored and the slowest took three minutes.

So I wrote down the decisions:

  • Hosted stays primary for anything that changes data or runs a command.
  • A local model gets writes only after zero critical failures, mutation safety of 1.00, and tool and argument accuracy of at least 0.98, on both a dev set and an untouched holdout.
  • More VRAM isn’t justified. A bigger card would not fix these failures.

The hosted model had one critical failure too, so the deterministic controls stay on for every model. I also exported 680 replayable cases from the agent’s history to grow the benchmark. They are blocked from training. Real household messages stay out of model weights.

A fine-tune that didn’t earn its place

The obvious next step was QLoRA on gpt-oss-20b. Training used 102 synthetic cases, 77 for training and 25 for validation. No household messages, no hosted-model outputs. The targets were mostly the stock model’s own answers, kept only when they scored perfectly.

The coding agent doing the work caught a format bug before any GPU time. The pinned tokenizer rendered training targets that ended in <|return|> with no Harmony final channel. At inference the model is expected to answer in an explicit final message. Training on the first version would have taught it a shape it never uses.

After the fix, training took about four minutes. On a 55-case holdout:

  • The 4-bit base scored 0.57. The adapter scored 0.80. Critical failures fell from 25 to 12.
  • Reads the base got right every time dropped to 0.025, mostly from parser failures.
  • Mean latency went from 4.5 s to 16.7 s.

The stock gpt-oss:20b I already run in Ollama scored 0.83 on an earlier version of that holdout. That number isn’t strictly comparable. It used a different runtime and an older prompt that showed the model hint-like case names. So I can’t say stock won by three points. I can say the adapter didn’t clearly beat what I already had, it broke reads, and it was slower than its own base. I did not deploy it.

Beating your own training base is easy to mistake for progress. Compare against the model you would actually ship, on the same harness.

A retrieval problem wearing a fine-tuning costume

The second project, gonk-sft, had a narrower goal: a Q&A model for hyperlocal public information, like businesses, places, hours, and phone numbers. QLoRA on a 4-bit Qwen2.5 14B, with a 7B for comparison, served on an OpenAI-compatible endpoint.

The harness came first. 131 frozen questions: 93 DEV, which the agent could read, and 38 blind TEST, which it could never see. A deterministic grader failed any answer with a phone number, address, price, or URL that wasn’t in the source, any answer missing required facts, and any invented value where the source said unknown.

Then I let a coding agent run the loop overnight and into the next day: train, grade, read the DEV failures, fix the data, retrain. Each run was kept or thrown away by its TEST score.

Stage TEST (of 38) DEV (of 93)
Bare 14B, no adapter 1 4
First 813-row training set 7 15
Best after hyperparameter tuning 9 22
First run trained to call a lookup tool 20 60
Best lookup run, 9 hits per query 28 60

Facts-only training stalled at 9. Hyperparameters added two. Rows aimed at DEV failures made TEST worse. The 7B tied the 14B. Questions about places scored 0 of 8 on TEST.

In the morning I asked whether it would be easier to give the model a tool than to make it memorize. The agent agreed: “That is a retrieval problem wearing a finetune costume.”

Phase two trained the model to emit a lookup_facts call, ran a keyword search over the facts file, and answered from the hits. TEST went from 9 to 20 on the first run and to 28 after tuning the lookup. Place questions went from 0 of 8 to 7 of 8. Fabrications on TEST fell from 11 to 2.

After that, every knob lost or tied: lookup scoring, the training rows, the learning rate. A sweep over lookup size gave 22, 24, 27, 28, and 23 for 6 through 10 hits. The agent called it a local maximum and stopped launching runs.

Weights are still good for voice, refusals, and saying “not published.” They are bad at being a phone book.

One caveat. The loop picked winners by TEST score. The agent never saw TEST answers, but more than 40 selections against 38 questions still leak a little. I’d treat 28 as optimistic. The next version gets a fresh holdout.

What local is good for today

Small, low-risk jobs. I wrote a router that sends four kinds of work to a local qwen3.5:35b-a3b: classifying a message, summarizing short text, picking a route, and cleaning up a draft. With thinking off, its smoke cases returned in about 0.5 to 0.9 seconds after the model loaded.

Anything that looks like a key, token, or password never reaches the local model. Neither do oversized inputs. Both go to the primary model. Host commands, writes, purchases, and advice are off the table. The router is not wired into the live agent yet.

What I’d tell someone starting

  • Local is fast. Fast is not the same as right, and the misses cluster where it matters.
  • Keep a hosted fallback. Route by risk.
  • Compare a fine-tune against the stock model you already run, on the same harness.
  • If the model needs to know facts, give it a lookup. Train it to use the lookup well.
  • Build the grader and the blind set before the first training run. They are the only reason I trust any of these numbers.