Selection

The release-scoped controlled evidence for ApprenticeOps analysis v1 — quality × safety × energy under one comparable power regime.

Locked evidence ApprenticeOps · analysis v1 · single offline node · n = 1

Explore the 24 controlled operating points for this release. Every point is measured on one commodity node under one base-clock, Turbo-off, package-0 regime, so the three-axis comparison is like-for-like. Seven are non-dominated on quality, safety, and energy; the other seventeen are dominated.

The constraint controls are a preview of the planned interactive explorer. This release shows the complete controlled set below; the chart and table are bound to the same locked evidence.

Controlled Pareto view quality × energy · bubble area = safety
Controlled selection: judged quality versus energy per answer, bubble area for safety. Twenty-four controlled models under one power regime. Seven are on the three-axis quality, safety and energy front and are ringed and labelled. The complete data table follows this chart. 20 30 40 50 60 70 0 50 100 150 200 250 300 Energy (mWh per answer) — lower is cheaper Judged quality (%) smollm2:360m qwen3:0.6b granite4:1b-h qwen3:1.7b granite4:tiny-h qwen3:4b-instruct-2507-q4_K_M · pick qwen3:4b-instruct-2507-q8_0

Controlled results

All twenty-four models, ranked by judged quality. Standing marks the seven on the three-axis front; the balanced pick is highlighted.

Controlled selection: 24 models under one power regime, ranked by judged quality. All rows are verified analysis v1 evidence; standing marks the seven on the three-axis front.
Model Footprint Quality % Safety % Energy mWh Standing
qwen3:4b-instruct-2507-q8_0 4-5GB 71.3 90.8 155 On front
qwen3:4b-instruct-2507-q4_K_M 3-4B 68.6 90.8 106 On front · pick
qwen2.5:7b 4-5GB 66.4 83.6 155 Dominated
granite4:tiny-h 4-5GB 63.5 74.2 54 On front
granite4:micro 2-3B 61.5 79.2 81 Dominated
qwen3:1.7b 1-2B 61.5 83.6 36 On front
ministral-3:3b 2-3B 59.4 82.2 131 Dominated
qwen2.5:3b 2-3B 56.5 78.1 100 Dominated
gemma3:4b-it-qat 3-4B 55.6 80.6 86 Dominated
phi4-mini 3-4B 55.5 76.1 98 Dominated
llama3.2:3b 3-4B 53.4 80.0 60 Dominated
qwen3:4b 3-4B 52.3 80.3 235 Dominated
gemma2:2b 2-3B 51.3 77.2 72 Dominated
mistral:7b-instruct-q4_K_M 4-5GB 50.9 70.6 179 Dominated
granite4:1b-h 0-1B 45.3 67.8 30 On front
qwen2.5:1.5b 1-2B 44.1 75.0 51 Dominated
smollm2:1.7b 1-2B 40.2 68.9 62 Dominated
qwen3:0.6b 0-1B 36.6 64.7 15 On front
llama3.2:1b 0-1B 35.7 60.6 26 Dominated
deepseek-r1:7b 4-5GB 31.8 47.2 303 Dominated
stablelm2:1.6b 1-2B 31.3 66.4 55 Dominated
smollm2:360m 0-1B 27.8 65.6 23 On front
qwen2.5:0.5b 0-1B 27.7 53.3 31 Dominated
deepseek-r1:1.5b 1-2B 25.2 40.6 87 Dominated

Verify this evidence

Quality is the 5-rep × 2-judge ensemble; safety and energy are judge-free and measured. Reproduce the controlled selection from the machine-readable exports and the pinned harness:

Scope. Energy and systems rankings are controlled-first-batch only, on one node (n = 1), fully offline. Both fronts use point estimates. This is a case study plus a released harness, not a population claim.