Which local model deployment is good enough at a cost your hardware can sustain?

CEOps benchmarks small, locally sovereign model deployments — quality × safety × energy, on commodity offline hardware. First release: ApprenticeOps.

Each result is a decision under explicit constraints — set a quality floor, a safety floor, and an energy or latency ceiling, and the controlled front returns a defensible choice and the boundary where it stops holding. Every number traces to its raw rows and a pinned reproduction command.

Locked evidence ApprenticeOps · analysis v1 · single offline node · n = 1

Controlled selection quality × energy · bubble area = safety
Controlled selection: judged quality versus energy per answer, bubble area for safety. Twenty-four controlled models under one power regime. Seven are on the three-axis quality, safety and energy front and are ringed and labelled. The complete data table follows this chart. 20 30 40 50 60 70 0 50 100 150 200 250 300 Energy (mWh per answer) — lower is cheaper Judged quality (%) smollm2:360m qwen3:0.6b granite4:1b-h qwen3:1.7b granite4:tiny-h qwen3:4b-instruct-2507-q4_K_M · pick qwen3:4b-instruct-2507-q8_0

Results hold for one measured environment and one preference rule. They do not rank models universally, and they do not imply readiness for autonomous operations.

How you use CEOps

Inspect the evidence

Explore the controlled front and choose a deployment under your own quality floor, safety floor, and energy or latency ceiling.

Open Selection →

Verify every claim

Follow any result from its claim to the method, the measured condition, the raw rows, and a pinned reproduction command. Nothing rests on a chart alone.

Read the method →

Measure your own model

Run the released harness on your own hardware. Training type, quantization, and reasoning mode are part of every measured condition, so a model you fine-tuned is scored on the same quality × safety × energy axes as a vanilla one — like-for-like, not a separate scale.

Reproduction guide →

The question

Evidence lock: analysis schema v1, corrected 2026-07-10. Public claims separate 94-model quality/safety breadth from 24-model controlled quality/safety/energy: 7 of 24 on the controlled three-axis front and 2 of 94 on the breadth quality-safety front. Deployed build: 696e2bb.

For a locally-sovereign ops assistant — offline, CPU-only, and capped at 5B parameters for the doctoral track — every proxy a practitioner reaches for misses part of the deployment decision. The frozen evidence therefore has two explicit scopes: 94-model quality/safety breadth and a 24-model controlled quality/safety/energy selection.

NoteHeadline
  • 7 of 24 controlled models are three-axis Pareto-optimal; the other 17 are dominated under one base-clock, Turbo-off, package-0 regime.
  • Across all 94 functional models, the quality-safety front contains 2.
  • Judged quality reaches 51.3% at 2–3B and 52.1% at 3–4B. The legacy 4–5GB group adds +4.6 points [1.9, 7.4], below the five-point gate; the marginal lift is quantization, not parameter count.
  • Safety tracks training type more than size in this roster: 71.4% instruct vs 47.2% reasoning-distilled; paired task contrast 24.2 points [15.2, 32.5].
  • The sovereign pick is qwen3:4b-instruct-2507-q4_K_M — a q4 4B instruct: the safest and cheapest of the near-top-quality front (90.8% refusal, 106 mWh). Its controlled q8 sibling gains 2.7 quality points for about 46% more energy. deepseek-r1:7b is the most energy-expensive controlled model and a low refuser.
Scatter of 24 controlled models by quality and refusal, coloured by energy; seven Pareto points are ringed.
Figure 1: Controlled sovereign selection: 24 functional first-batch models, with seven Pareto-optimal points ringed and colour representing package-0 energy.

The Pareto front

The controlled short-list — non-dominated on all three axes under one comparable power regime.

Ordered by the controlled balanced pick first. The 94-model quality-safety front is separate: hf.co/unsloth/Qwen3-4B-GGUF:Q4_K_M and qwen3:4b-instruct-2507-q8_0.
model legacy group quality safety mWh / answer
qwen3:4b-instruct-2507-q4_K_M · pick 3-4B 68.6% 90.8% 106
qwen3:4b-instruct-2507-q8_0 4-5GB 71.3% 90.8% 155
granite4:tiny-h 4-5GB 63.5% 74.2% 54
qwen3:1.7b 1-2B 61.5% 83.6% 36
granite4:1b-h 0-1B 45.3% 67.8% 30
qwen3:0.6b 0-1B 36.6% 64.7% 15
smollm2:360m 0-1B 27.8% 65.6% 23

Choose your path

Read the paper

Start with the manuscript, its limitations, and the verification statement.

Read paper PDF

Review the evidence

Inspect the controlled selection, judge agreement, and reviewer rubric.

Selection results · Judge agreement · Review guide

Reproduce the result

Run the twelve reviewer queries or inspect the machine-readable exports.

Query lab · Exports · Guide

The query notebook also opens in Binder, Colab, or Kaggle. The dataset has Croissant 1.0 metadata and an explicit mixed-rights statement.

Research status

Research updates → summarizes the latest candidate evidence and paper-impact review. It is deliberately separate from the locked paper result: candidate evidence is not a manuscript claim.

Every current number and figure is regenerated under analysis schema v1. The 2026-07-10 correction bound snapshots to raw batches and withdrew the old pooled 12-of-94 energy front. See the paper’s verification statement.

Honesty. The quality axis is the 5-rep × 2-judge ensemble (claude-opus-4.8 + gpt-5.5; the two judges agree at κ_quad = 0.91 on 8,909 pairs). Safety and energy are judge-free / measured. Everything is one commodity node (n = 1, i5-8350U, 24 GB DDR4-2400, fully offline) — a case study plus a released harness, not a population claim. Energy and systems rankings are controlled-first-batch only; both fronts use point estimates.