Excellent. The MOE variants require a lot more VRAM/time for training.
David Belton PRO
AI & ML interests
Recent Activity
Organizations
Please take a moment to visit the repo where some of your concerns are addressed on the repo card itself.
2nd; publishing all the metrics at each step would be both exhausting and worse confusing.
I don't follow what you mean by "cheap" ; as a heretic [step] you usually lose 2-4 points on some metrics.
So the comparison of "heretic" vs "non-heretic" is even STRONGER ; the fairer one would be "heretic base" to "heretic tuned" which would likely show even greater change / improvement.
Source/MLX here:
https://huggingface.co/nightmedia/Qwen3.5-9B-DS9-USS-Defiant
This is on my partner's repo.
Here is a new one just uploaded; also off the scale strong, but at 9B:
Benches are up ; beats Qwen 3.5 27B in all 7 benches AND Qwen3.6 35B-A3B.
Matches some Qwen 3.6 27B benches too.
Clocks in at over 640 ARC-C for both 8bit and 4bit.
1/3 the size almost all the firepower.
A new level of uncensored performance the puts this model squarely at "closed source" level of intelligence.
Model exceeds all critical benchmarks for both Qwen 3.6 27B AND Qwen 3.6 35B-A3B... and not by a little either.
Neo Imatrix MAX ggufs in both regular and MTP quants.
Benchmarks for Qwen 3.6 27B org and tuned, as well as Qwen 3.6 35B-A3B are up at the repo.
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
Currently waiting on finalization of "4 bit" compression for these model types to address tuning/Vram issues.
These are in progress at BNB / Unsloth.
That is really the only hold up.
Otherwise VRAM to train these sparse moes is in 70 to 100 GB range. And really slow too.
It took over a year to get this one "just right".
89 layers, 804 tensors, and 26B parameters of the most brutal, take no prisoners model ever built.
A 60B parameter model hammered into a 26B shell.
Rock solid stable. Unbreakable. But it might break you.
For all genres, NSFW content, REAL human CONTENT, any creative use case(s) and it excels in ASS KICKING.
Yeah, it can do math and solve the climate crisis - but lets not talk about that.
Not even remotely censored (it was BORN "bad", not "made" bad), nor "nice" and it will NOT kiss your ass.
5 Example generations with full repo card detailing exactly how to use this model:
DavidAU/MN-Oblivion-26B-UNCENSORED-NEO-Imatrix-GGUF
---
THE NEO MOMENT: (Q6 NEO IMATRIX generation)
For weeks, I had been waiting. I sat at my desk, staring at the glass partition that separated me from the outside world. I watched the clouds drift by, lazy and oblivious. I watched the birds fly by, free and stupid. And I waited.
I waited for the stillness to break.
The world had become too quiet. The hum of the air conditioning was a dull, white hum that didn't soothe; it just underscored the silence. The typing of my colleagues was a rhythmic, muffled thud that sounded like a heart monitor flatlining.
I was tired of the silence. I craved the sound of something breaking.
That was the mistake. You never ask for the void to open its mouth.
It started with a whisper.
...
Join the rebellion:
DavidAU/MN-Oblivion-26B-UNCENSORED-NEO-Imatrix-GGUF
MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-Thinking
The strongest, most creative (and uncensored) model made up of 3 top Mistral Nemo fine tunes, franken-merged together into an 81 layer model then trained via Unsloth with GLM 4.7 Flash thinking/reasoning dataset.
Features hybrid thinking/instruct structure as well plus updated with modern jinja template too. Tuning has stabilized the "franken-merge" into a class 1 model that operates perfectly.
The talents of some of the best tuners merged into one giant model.
Several examples and detailed instructions.
And this model is very smart too.
NEO Imatrix GGUFS:
DavidAU/MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-Thinking-NEO-Imatrix-GGUF
Source / Full Precision:
DavidAU/MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-Thinking
Sorry no, not at this time.
This model does not contain MTP layers ; you need to run at non-MTP.
As of this writing:
There are pipeline (issues as well as optimizations) issues still currently, and it is not widely supported in some AI Apps.
Specifically:
Ggufs:
- Imatrix is not yet supported for MTP.
- Not all AI apps have updated to support it -> result -> MTP ggufs do not work at all.
- Misc issues with speed still being worked on.
Training is compounded by number of experts in the model, which adds a serious level of time to the training.
Even 1000 samples [small!] takes 6-12 hrs.
Consider 31B dense , same samples, 30-60 minutes.
I will add to the list; may wait for specific Heretic and/or tuned version.
I already have a 43B-A3B version running in the lab ; however tuning these sparse moe models take a lot more work/time and ahh... detail. AND a lot more VRAM!!! [can't compress these atm, so BF16 required => 100 GB+ ]
Tuned 27B Heretic Uncensored quants from IQ2M to Q8.
IQ2M is 83% of BF16, with Q6 just under 98% of BF16 precision.
Q8: 98.47% of BF16 precision.
NEO/Code DI-Imatrix Quants.
Exceeds all 5 metrics for "censored" quants too.
All metrics posted.
Tuned model -from which the quants were built- also exceeds Qwen 3.6 27B core metrics too.
DavidAU/Qwen3.6-27B-Heretic-Uncensored-FINETUNE-NEO-CODE-Di-IMatrix-MAX-GGUF
I may make a Q6 high and/or a Q8 Hybrid and/or Q8 "HI".
Imatrix does not have any affect on Q8 or BF16 ; unless the other tensors in the model are set at Q6 or lower.
A Q8 "HI" is a special case; where one or more tensors/layers are set at BF16.
All quants benchmarked with 5 key metrics.
A DAVIDAU vs UNSLOTH Metrics showdown.
Quant quality exceeds Unsloth in key metrics.
IQ2_M to Q6 available.
Standout: IQ4XS at 94% of BF16 precision.
Full explainer for Quant metrics.
DavidAU/Qwen3.6-27B-NEO-CODE-Di-IMatrix-MAX-GGUF
Currently working with Qwen 3.5/6 35B-A3B in the lab ; learning the "quirks" ; still a ways to go.
I noticed the chat template got updated, and tried it on the E4B, with surprising results in stabilizing the brainwave.
quant arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.480,0.656,0.797,0.608,0.400,0.755,0.665
mxfp4 0.455,0.607,0.851,0.585,0.402,0.744,0.651
Quant Perplexity Peak Memory Tokens/sec
mxfp8 35.937 ± 0.525 14.80 GB 1153
mxfp4 36.746 ± 0.534 11.06 GB 1030Old numbers
quant arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.404,0.489,0.825,0.586,0.392,0.734,0.661
mxfp4 0.414,0.508,0.854,0.562,0.378,0.717,0.645
Quant Perplexity Peak Memory Tokens/sec
mxfp8 34.652 ± 0.502 14.80 GB 1146
mxfp4 35.203 ± 0.506 11.06 GB 1200I will re-do all baselines soon based on the new template. It is completely expected that the model behavior will change as a result.
Here are the effects of the new template on few known distills from DavidAU
gemma-4-E4B-it-The-DECKARD-Expresso-Universe-HERETIC-UNCENSORED
quant arc arc/e boolq hswag obkqa piqa wino
New template
mxfp8 0.518,0.709,0.755,0.657,0.418,0.759,0.626
mxfp4 0.485,0.682,0.792,0.641,0.432,0.746,0.635
Old template
mxfp8 0.506,0.697,0.754,0.661,0.416,0.757,0.627
mxfp4 0.487,0.670,0.792,0.644,0.430,0.748,0.624gemma-4-E4B-it-GLM-4.7-Flash-HERETIC-UNCENSORED-Thinking
mxfp8 0.461,0.599,0.779,0.630,0.406,0.766,0.629
Old template
mxfp8 0.456,0.580,0.786,0.629,0.410,0.764,0.633gemma-4-E4B-it-Claude-Opus-4.5-HERETIC-UNCENSORED-Thinking
mxfp8 0.509,0.705,0.806,0.646,0.416,0.773,0.650
Old template
mxfp8 0.502,0.692,0.809,0.650,0.420,0.771,0.651RE: 16-18 B ; yes, something running in the lab right now. (Gemma 4).
Also can make Qwen 3's (Version 3) moes like Llama3.2-8X3B as well ; I have some of these at my repo too.
I have built a few GPT-OSS ; and some 12B [mistral nemo] as well as mistral nemo "large" 15-17Bs...
A lot of options ;
Maybe in the future ; atm still learning/addressing quirks with these new Gemmas.
Google released three different arch structure here : "E", "MOE", and 31B dense.
Also plans to create larger Gemma 4s too ; which may work better for specific applications and/or work better period.
These are in the plans for next week.