grug-27b-gguf

2026-07-23: all rocks re-squeezed from v2.1 weights (deep think on hard problems, stuck-loop escape, stop discipline - full changelog on grug-27b card). re-download if you grab rocks before. mmproj unchanged (vision tower untouched).

grug brain squeezed into small rock. run on your cave computer with llama.cpp.

this GGUF of grug-27b: Qwen3.6-27B that think in dense grug-speak inside <think>, answer in normal english. same reasoning depth, way fewer think token. full story on main model card.

27b and 35b hunt same prey

both parent grug hunt HumanEval and sanitized MBPP. number below come from big parent brain, NOT squeezed GGUF rock. grug not claim rock test it never get. number show pass@1 percent. bold grug win that hunt.

hunt grug-27b v2.1 grug-35b rebuilt
HumanEval (164) 87.2 80.5
MBPP sanitized (100) 85.0 88.0

rock sizes

file quant size grug opinion
grug-27b-Q8_0.gguf Q8_0 28.6 GB basically bf16. big rock.
grug-27b-Q6_K.gguf Q6_K 22.1 GB very good rock
grug-27b-Q5_K_M.gguf Q5_K_M 19.2 GB good rock
grug-27b-Q4_K_M.gguf Q4_K_M 16.5 GB best size/smart trade. grug pick this.
grug-27b-Q3_K_M.gguf Q3_K_M 13.3 GB small rock. smart mostly survive.
mmproj-grug-27b-f16.gguf mmproj f16 see repo eye rock. give grug vision back.

every rock load-tested with llama.cpp before upload. no missing-tensor sickness (grug check twice now, learn from 9b).

Q4 person? special rock exist

grug make QAT version of Q4_K_M: weights trained while feeling 4-bit rounding rock before final squish. better Q4 quality, same grug brain: grug-27b-qat-q4-gguf. rocks here best for Q8/Q6/Q5 people.

if rock act broken

single-token spam ("/" forever etc) = NOT the rock. hybrid DeltaNet brain CANNOT survive llama.cpp context-shift: old builds shift on context overflow and corrupt the recurrent state into token spam. fix:

  • use RECENT llama.cpp (qwen3_5 support; new builds refuse instead of shift)
  • agent frontends (OpenCode etc): set -c 16384 or bigger
  • still broken? re-download rock (verify size) + check backend grug re-test rock after every report: loads clean, zero spam at proper config.

how run

need recent llama.cpp (qwen3_5 arch support).

llama-server -m grug-27b-Q4_K_M.gguf -c 16384 --temp 0.6 --top-p 0.95 --top-k 20
  • vision NOW work: pair any quant with mmproj-grug-27b-f16.gguf (llama-server -m grug-27b-Q4_K_M.gguf --mmproj mmproj-grug-27b-f16.gguf). MTP still not included.
  • context: base support 262144, pick what your RAM allow
  • thinking on by default, reasoning arrive inside <think>...</think>
  • for agent frameworks (OpenCode etc): works with think-stripped history, grug trained for exactly that world

grug made by ProCreations. base brain by Qwen team.

Downloads last month
2,108
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ProCreations/grug-27b-gguf

Base model

Qwen/Qwen3.6-27B
Quantized
(8)
this model