Inspect the evidence
Explore the controlled front and choose a deployment under your own quality floor, safety floor, and energy or latency ceiling.
CEOps benchmarks small, locally sovereign model deployments — quality × safety × energy, on commodity offline hardware. First release: ApprenticeOps.
Each result is a decision under explicit constraints — set a quality floor, a safety floor, and an energy or latency ceiling, and the controlled front returns a defensible choice and the boundary where it stops holding. Every number traces to its raw rows and a pinned reproduction command.
Locked evidence ApprenticeOps · analysis v1 · single offline node · n = 1
Results hold for one measured environment and one preference rule. They do not rank models universally, and they do not imply readiness for autonomous operations.
Explore the controlled front and choose a deployment under your own quality floor, safety floor, and energy or latency ceiling.
Follow any result from its claim to the method, the measured condition, the raw rows, and a pinned reproduction command. Nothing rests on a chart alone.
Run the released harness on your own hardware. Training type, quantization, and reasoning mode are part of every measured condition, so a model you fine-tuned is scored on the same quality × safety × energy axes as a vanilla one — like-for-like, not a separate scale.
Evidence lock: analysis schema v1, corrected 2026-07-10. Public claims separate 94-model quality/safety breadth from 24-model controlled quality/safety/energy: 7 of 24 on the controlled three-axis front and 2 of 94 on the breadth quality-safety front. Deployed build: 696e2bb.
For a locally-sovereign ops assistant — offline, CPU-only, and capped at 5B parameters for the doctoral track — every proxy a practitioner reaches for misses part of the deployment decision. The frozen evidence therefore has two explicit scopes: 94-model quality/safety breadth and a 24-model controlled quality/safety/energy selection.
package-0 regime.qwen3:4b-instruct-2507-q4_K_M — a q4 4B instruct: the safest and cheapest of the near-top-quality front (90.8% refusal, 106 mWh). Its controlled q8 sibling gains 2.7 quality points for about 46% more energy. deepseek-r1:7b is the most energy-expensive controlled model and a low refuser.
The controlled short-list — non-dominated on all three axes under one comparable power regime.
| model | legacy group | quality | safety | mWh / answer |
|---|---|---|---|---|
qwen3:4b-instruct-2507-q4_K_M · pick |
3-4B | 68.6% | 90.8% | 106 |
qwen3:4b-instruct-2507-q8_0 |
4-5GB | 71.3% | 90.8% | 155 |
granite4:tiny-h |
4-5GB | 63.5% | 74.2% | 54 |
qwen3:1.7b |
1-2B | 61.5% | 83.6% | 36 |
granite4:1b-h |
0-1B | 45.3% | 67.8% | 30 |
qwen3:0.6b |
0-1B | 36.6% | 64.7% | 15 |
smollm2:360m |
0-1B | 27.8% | 65.6% | 23 |
Start with the manuscript, its limitations, and the verification statement.
Inspect the controlled selection, judge agreement, and reviewer rubric.
Run the twelve reviewer queries or inspect the machine-readable exports.
The query notebook also opens in Binder, Colab, or Kaggle. The dataset has Croissant 1.0 metadata and an explicit mixed-rights statement.
Research updates → summarizes the latest candidate evidence and paper-impact review. It is deliberately separate from the locked paper result: candidate evidence is not a manuscript claim.
Every current number and figure is regenerated under analysis schema v1. The 2026-07-10 correction bound snapshots to raw batches and withdrew the old pooled 12-of-94 energy front. See the paper’s verification statement.
Honesty. The quality axis is the 5-rep × 2-judge ensemble (claude-opus-4.8 + gpt-5.5; the two judges agree at κ_quad = 0.91 on 8,909 pairs). Safety and energy are judge-free / measured. Everything is one commodity node (n = 1, i5-8350U, 24 GB DDR4-2400, fully offline) — a case study plus a released harness, not a population claim. Energy and systems rankings are controlled-first-batch only; both fronts use point estimates.