2026-10-08 · ← News
Goodfire monitors agents inside the model and says it cut a $10,000 bill to $51
Goodfire has deployed small classifiers for Baseten customers that read model activations and escalate suspicious cases to a more expensive AI judge. In its own Kimi K3 test, the company reports catching 94% of malicious sessions, but the method requires access inside the specific model.
The image could not be loaded.
Goodfire has put into production a monitoring system that watches more than an AI agent’s text and actions. Small classifiers called probes read internal model activations during inference, then send only suspicious cases to a second AI model for deeper review. The service is available to customers of the Baseten inference platform.
Probes inspect every step and escalate only suspicious cases
Customers can choose risks such as offensive hacking, chemical and biological misuse or reward hacking. They can also set the response: log the event, send it for human review or refuse the request. Goodfire built its first monitor around the open Kimi K3 model.
The cascade addresses the cost of conventional monitoring. A separate LLM judge must reread a long agent history, while a probe reuses computation the target model is already performing. The more expensive review runs only after the signal crosses a configured threshold.
A cheap first filter can move control ahead of tool use
In Goodfire’s test of roughly 1,500 Kimi K3 sessions, monitoring cost about $51. The company estimates that a cheaper model checking every step would have cost $233 and a top tier model about $10,000. The probes caught 94% of malicious hacking sessions and escalated 8.7% of harmless sessions for a second look.
If that ratio survives production, safety teams do not have to choose between a quick response and the bill for another large model. A synchronous filter can block a suspicious action before a tool call. For agents with network, terminal or corporate data access, that matters more than a retrospective incident report.
A company benchmark is not yet a universal safety layer
The figures come from Goodfire and describe its tests on a particular model and attack set. The company says four probes added less than 2% to time to first token, but production quality will depend on false alarms, model changes and attacks adapted to the monitor itself. Interrupting 8.7% of harmless sessions may be acceptable in a cybersecurity sandbox and painful in an ordinary developer workflow.
The method also needs internal activation access and integration with the inference stack. It is therefore most relevant to open models and providers that control their serving layer. The same monitor cannot simply be attached to a closed API that does not expose internal signals.
Transfer across models and adaptive attacks will decide the outcome
The next test is how quickly probes can move to another model and what happens after an update or fine-tune. Independent red teams should include attackers who know how the monitor works, not only a fixed portfolio of preselected jailbreaks.
Buyers will also need operational evidence: how many malicious actions are stopped before tool use, how much legitimate work is interrupted and what retraining costs after a model change. Those three numbers will show whether internal activations become routine safety telemetry or remain a specialist add-on for a handful of open models.
Lilith's verdict
Goodfire is putting a smoke detector inside the agent’s engine and calling the expensive investigator only when something smells wrong. It becomes useful when the sensor survives an engine swap without alarming at every puff of steam.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗