BeefAndPoultry@lemmus.org to LocalLLaMA@sh.itjust.worksEnglish · 9 days agounsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Facehuggingface.coexternal-linkmessage-square7linkfedilinkarrow-up124arrow-down11file-text
arrow-up123arrow-down1external-linkunsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Facehuggingface.coBeefAndPoultry@lemmus.org to LocalLLaMA@sh.itjust.worksEnglish · 9 days agomessage-square7linkfedilinkfile-text
minus-squareMultiplexer@discuss.tchncs.delinkfedilinkEnglisharrow-up3·9 days agoAnyone knows, why 4bit quant is only marginally smaller than 8bit quant, though? And shouldn’t 8bit be roughly equal to the number of parameters anyway, so ~300GB?
minus-squareBeefAndPoultry@lemmus.orgOPlinkfedilinkEnglisharrow-up4·9 days agothe original model has a lot of parts that were natively trained in 4 bit, so those layers can’t go higher
Anyone knows, why 4bit quant is only marginally smaller than 8bit quant, though?
And shouldn’t 8bit be roughly equal to the number of parameters anyway, so ~300GB?
the original model has a lot of parts that were natively trained in 4 bit, so those layers can’t go higher