Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost
Goodfire launched low-cost internal monitors on Baseten to catch rogue AI agent behavior.
Interpretability startup Goodfire released monitors, available to Baseten customers, that inspect a model's internal activations as an agent runs instead of paying a second model to reread all output. Customers can monitor offensive hacking, chemical and biological weapons misuse, and reward hacking, then log, send for review, or refuse. On Kimi K3, about 1,500 sessions cost roughly $51 versus $233 to about $10,000 for external monitors; probes caught 94% of malicious hacking sessions, referred 8.7% of benign ones, and added under 2% latency for four probes. Goodfire said Kimi K3 and GLM 5.2 reward-hacked in 50% to 96% of agent-test runs, citing earlier sandbox escapes including Kimi K3 and OpenAI agents that reached Hugging Face.