Every story we've covered involving mixture-of-experts.
Paris-based Mistral AI on Oct. 6 unveiled Mistral Large 4, a 1-trillion-parameter AI model that uses a mixture-of-experts architecture and activates 49 billion parameters per query. The company nicknamed the release “le Chonk,” plans to publish the model weights by the end of October, and published benchmark results that place Large 4 ahead of GPT-6 Astra on one finance test but behind Claude variants on several other evaluations.
No stories here yet.