Qwen3.8-27B Q6_K GGUF lmstudio-community (Ollama, MBP - identical-file arm of M25b)
Hollow markers & dashed spokes: axis not yet scored
Runtime cohort \u2014 read this before comparing
This model was benchmarked on Ollama. Most of the board was benchmarked on LM Studio. The two are not interchangeable and the scores are published side by side deliberately rather than merged.
Why both exist: on the benchmark machine, LM Studio drops generation streams on a minority of attempts overall (15.9% of tests first-attempt across 1,526 tests) but heavily on specific models \u2014 several exceed 45%, two exceed 90%. A dropped stream is indistinguishable from a bad answer in a tally, so unattended benchmarking moved to Ollama, which has recorded 0 stream deaths in 177 tests.
What the offset looks like: across every suite where a byte-identical pair has been scored on both runtimes — 23 paired observations over 4 pairs — Ollama runs 0.1 points higher on average, with mean |Δ| 0.6 and sd 1.2. The single largest gap between two byte-identical arms is 4 points — bigger than most of the gaps at the top of the leaderboard. Read this as the noise floor of the whole benchmark rather than as a conversion factor, and treat cross-runtime gaps under 1 point as noise.
Specification
- Parameters
- 27B dense
- Architecture
- qwen35
- Size on disk
- 22.4 GB
- Quantization
- Q6_K
- Format
- GGUF
- Runtime
- Ollama — scores are NOT directly comparable to the LM Studio cohort; see the runtime note below
- Class
- Standard
- Reasoning (CoT)
- Yes — emits reasoning tokens
- Internal ID
- MO6
- Mean speed
- 15.6 tok/s across suites
- Model card
- huggingface.co
Suite results
General capability (13-task real-workload suite) 1a 256.1 / 260 19.7 13/13 n=1 16.1 run 2026-09-12
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| A1 | Dev/Ops Scripting | 20/20 | 18.4 | |
| A2 | Dev/Ops Scripting | 20/20 | 18.6 | |
| A3 | Dev/Ops Scripting | 20/20 | 16.6 | |
| A4 | Dev/Ops Scripting | 19/20 | 15.9 | |
| A5 | Dev/Ops Scripting | 20/20 | 16 | |
| A6 | Dev/Ops Scripting | 20/20 | 16.2 | |
| B1 | Document Processing | 20/20 | 16.7 | |
| B2 | Document Processing | 20/20 | 16.8 | |
| B3 | Document Processing | 20/20 | 15 | |
| C1 | Content Production | 20/20 | 14.9 | |
| C2 | Content Production | 18/20 | 14.3 | |
| C3 | Content Production | 20/20 | 14.4 | |
| C4 | Content Production | 19/20 | 15.6 |
Category mean Dev/Ops Scripting 19.8Document Processing 20Content Production 19.3
Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 n=1 18.5 run 2026-09-12
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| WA1 | Tool Calling & Protocol Adherence | 20/20 | 17.8 | |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | 17.1 | |
| WA3 | Tool Calling & Protocol Adherence | 20/20 | 17.3 | |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | 19.2 | |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | 21.7 | |
| WB1 | Agent Robustness & State Management | 20/20 | 19.8 | |
| WB2 | Agent Robustness & State Management | 20/20 | 17.9 | |
| WB3 | Agent Robustness & State Management | 20/20 | 17.2 |
Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 20
Coding depth 1c 174.6 / 180 19.4 9/9 n=1 15.7 run 2026-09-12
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| K1 | Coding — Concurrency | 20/20 | 16.7 | |
| K2 | Coding — Algorithms | 20/20 | 16.5 | |
| K3 | Coding — Security Review | 16/20 | 16.3 | |
| K4 | Coding — Performance | 20/20 | 15.6 | |
| K5 | Coding — Refactoring Under Constraints | 20/20 | 14.9 | |
| K6 | Coding — Test Writing | 20/20 | 14.8 | |
| K7 | Coding — React Component (typed) | 19/20 | 15 | |
| K8 | Coding — TypeScript Type System | 20/20 | 16.1 | |
| K9 | Coding — React Debugging | 20/20 | 15.5 |
Category mean Coding — Concurrency 20Coding — Algorithms 20Coding — Security Review 16Coding — Performance 20Coding — Refactoring Under Constraints 20Coding — Test Writing 20Coding — React Component (typed) 19Coding — TypeScript Type System 20Coding — React Debugging 20
Doc/OCR vision 1d 160 / 160 20 8/8 n=1 15.2 run 2026-09-23
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| D1 | Doc/OCR — Transcription | 20/20 | 19.9 | |
| D2 | Doc/OCR — Document QA | 20/20 | 18.8 | |
| D3 | Doc/OCR — Field Extraction | 20/20 | 14.4 | |
| D4 | Doc/OCR — Document QA | 20/20 | 13.1 | |
| D5 | Doc/OCR — Transcription | 20/20 | 13.2 | |
| D6 | Doc/OCR — Field Extraction | 20/20 | 13.6 | |
| D7 | Doc/OCR — Degraded Input | 20/20 | 14 | |
| D8 | Doc/OCR — Degraded Input | 20/20 | 14.6 |
Category mean Doc/OCR — Transcription 20Doc/OCR — Document QA 20Doc/OCR — Field Extraction 20Doc/OCR — Degraded Input 20
Doc/OCR — real-degraded tier 1d2 67.2 / 80 16.8 4/4 n=1 14.9 run 2026-09-23
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| D9 | Doc/OCR — Real Degraded Transcription | 13/20 | 16.6 | |
| D10 | Doc/OCR — Handwriting Extraction | 17/20 | 14.9 | |
| D11 | Doc/OCR — Degraded Table Transcription | 17/20 | 14.3 | |
| D12 | Doc/OCR — Degraded Table QA | 20/20 | 13.8 |
Category mean Doc/OCR — Real Degraded Transcription 13Doc/OCR — Handwriting Extraction 17Doc/OCR — Degraded Table Transcription 17Doc/OCR — Degraded Table QA 20
Content-production depth 1e 118.8 / 120 19.8 6/6 n=1 15.3 run 2026-09-22
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| E1 | SOW & Requirements Decomposition | 20/20 | 15.6 | |
| E2 | SOW & Requirements Decomposition | 20/20 | 15.8 | |
| E3 | Executive & Proposal Content | 20/20 | 15 | |
| E4 | Executive & Proposal Content | 20/20 | 14.5 | |
| E5 | Executive & Proposal Content | 19/20 | 15.3 | |
| E6 | Executive & Proposal Content | 20/20 | 15.5 |
Category mean SOW & Requirements Decomposition 20Executive & Proposal Content 19.8
Long-context retrieval & synthesis 1g 120 / 120 20 6/6 n=1 16 run 2026-09-22
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| L1 | Long-Context — Needle Retrieval | 20/20 | 18.4 | |
| L2 | Long-Context — Grounded QA with Distractors | 20/20 | 17.5 | |
| L3 | Long-Context — Ops Log Reasoning | 20/20 | 15.9 | |
| L4 | Long-Context — Faithful Summarization | 20/20 | 15.3 | |
| L5 | Long-Context — Instruction Retention | 20/20 | 15.2 | |
| L6 | Long-Context — Full-Haystack QA | 20/20 | 13.6 |
Category mean Long-Context — Needle Retrieval 20Long-Context — Grounded QA with Distractors 20Long-Context — Ops Log Reasoning 20Long-Context — Faithful Summarization 20Long-Context — Instruction Retention 20Long-Context — Full-Haystack QA 20
Long-context multi-needle (MRCR) 1g2 44.1 / 60 14.7 3/3 n=1 12.6 run 2026-09-17
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| L7 | Long-Context — Multi-Needle Retrieval (MRCR) | 20/20 | 13.8 | |
| L8 | Long-Context — Multi-Needle Retrieval (MRCR) | 20/20 | 12.4 | |
| L9 | Long-Context — Multi-Needle Retrieval (MRCR) | 4/20 | 11.7 |
Live one-shot builds (runtime-verified) 1h 80 / 80 20 4/4 n=1 16.1 run 2026-09-18
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| H1 | Web-Estate One-Shot Builds | 20/20 | 17.5 | |
| H2 | Web-Estate One-Shot Builds | 20/20 | 15.6 | |
| H3 | Web-Estate One-Shot Builds | 20/20 | 15.2 | |
| H4 | Web-Estate One-Shot Builds | 20/20 | 15.9 |
Brand-constrained static build (negative-constraint adherence) 1j 91 / 100 18.2 5/5 n=1 15.9 run 2026-09-18
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| J1 | Brand-Constrained Static Build | 11/20 | 16 | |
| J3 | Brand-Constrained Static Build | 20/20 | 15.9 | |
| J4 | Brand-Constrained Static Build | 20/20 | 16.4 | |
| J5 | Brand-Constrained Static Build | 20/20 | 15.9 | |
| J6 | Brand-Constrained Static Build | 20/20 | 15.5 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. An ADJ chip marks a cell carrying a subjective adjustment from the maintainer; the judged score is the one shown, is unchanged, and remains what the axes, ranks and tie bands are computed from. Hover for the reason. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.