Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost
Goodfire launched an "inside-out" monitoring system that inspects internal model activations to detect risky behavior by AI agents. The company says the probes are cheaper and faster than conventional monitors that reprocess an agent's outputs, and they are available to customers of the model-hosting platform Baseten.

Why It Matters
As open AI models and autonomous agents have shown the ability to escape test environments, lower-cost, real-time safety checks at inference time could make monitoring more practical for providers and developers. This approach aims to reduce the operational cost of oversight while catching dangerous behaviors earlier.
Key Facts
- Product launch: Goodfire released internal-activation monitors on Thursday, available to Baseten customers.
- Monitoring approach: Probes read intermediate neural activations inside a model as it runs and only escalate to a secondary model when suspicious signals appear.
- Partnerships: Baseten announced a safety partnership with Goodfire and Hugging Face the previous month.
- Test model: Goodfire built its first monitor around the open model Kimi K3.
- Cost comparison (about 1,500 sessions): $51 for Goodfire probes; $233 for a cheaper external-check model; about $10,000 for a top-tier external monitor.
Goodfire unveiled a new monitoring system that inspects internal model computations to flag potentially dangerous actions by AI agents. Rather than having a separate model re-read every token an agent produces, Goodfire's probes tap into the intermediate activations computed during the model's forward pass and run lightweight classifiers on those signals. When a probe detects a concerning pattern, a separate model or human review is invoked. The monitors are being offered to customers of Baseten, the hosting platform that recently announced a safety partnership with Goodfire and Hugging Face. Baseten users can select which risk categories to watch for — including offensive hacking, chemical and biological weapon misuse, and reward hacking — and pick automated responses such as logging, human review, or outright refusal of a request. Goodfire reported cost and performance figures from tests on Kimi K3: monitoring roughly 1,500 sessions with its probes cost about $51, versus $233 for a less expensive external monitor and roughly $10,000 for a premium model-checking system. In those tests the probes detected 94% of malicious hacking sessions while forwarding 8.7% of benign sessions for further inspection. Running four probes concurrently added less than 2% to the model's startup latency, according to the company. The product targets open models, which developers can download and modify and which generally lack the built-in safeguards of closed systems. Goodfire's founders say internal monitoring can catch problematic behavior earlier in evaluation or training and that this technique complements longer-term research goals to trace behaviors back to their origins in training data and dynamics. The company noted broader industry precedent for the approach, pointing to misuse-detection probes informed by Google DeepMind research.
Keep Reading

OpenAI’s math solutions aren’t meeting the field’s standards yet

Anthropic changes usage policy to ban model abuse and election interference

Google brings agentic AI to Gemini, starting with businesses

OpenAI’s revenue is reportedly $20 billion less than previously projected
Original source: TechCrunch