Johannes Uusikuu
Johneeee
AI & ML interests
Small omlx models without vision & mtp for m1 m2 or low memory macs. Try the last4native or last4_8bit quants they have 5 bit base but bump a lot of the model to 8 and 6 bit. The last 4 layers effect the output quality insurprising ways. My experimentation shows that giving a big bit budget bumps up 24 % of the tensors to 8 bit and around 15 percent to 6 bit. I do also experimentation with the last 4 layers in either 8 bit or bf16.
Recent Activity
new activity about 3 hours ago
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP:GGUF special tensor bump quantifizations? new activity about 7 hours ago
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP:PrismAura quantized version (DGX Spark optimized) new activity about 7 hours ago
trithemius/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-MTP-PrismAura-5.5bit:Doing MLX quant based on thisOrganizations
None yet