Every story we've covered involving model-interpretability.
Baseten and its research arm Base Labs announced a new safety infrastructure standard for open-weight AI models in partnership with Hugging Face and Goodfire AI. The initiative aims to build evaluation and monitoring tools for open models amid growing concerns about abliterated models that remove safety controls.
No stories here yet.