Every story we've covered involving open-models.
Goodfire launched an "inside-out" monitoring system that inspects internal model activations to detect risky behavior by AI agents. The company says the probes are cheaper and faster than conventional monitors that reprocess an agent's outputs, and they are available to customers of the model-hosting platform Baseten.
No stories here yet.