
The Language of Surrender: Larger Model Measurements
Three models one size class up — Qwen3-14B, Phi-4, and Mistral-Small-24B at Q4_K_M quantization — measured on the same 1,505-sentence limitation gold set, same prompt, same harness as the earlier six-model comparison. Qwen3-14B scores F1 0.8272 as a single model, above the previous two-model union (0.8143) and the published rule-system reference (0.800). VRAM at load: 10.6 GB of 16.3 GB. Every number is a full-split measurement.
