Every story we've covered involving smarter.
In brief Multiverse Computing's team published a method called Quantization-Aware Healing on the Hugging Face blog on August 25. They shrank OpenAI's open GPT-OSS model from 120 billion parameters to 60 billion and compressed its memory to 4-bit—and the small version beat the full-quality model it was copied from on 7 of 9 tests.
No stories here yet.