Lilith.
⌕
Editorial illustration: Goodfire monitors agents inside the model and says it cut a $10,000 bill to $51
Lilith illustration · editorial remix

Goodfire has put into production a monitoring system that watches more than an AI agent’s text and actions. Small classifiers called probes read internal model activations during inference, then send only suspicious cases to a second AI model for deeper review. The service is available to customers of the Baseten inference platform.

Probes inspect every step and escalate only suspicious cases

Customers can choose risks such as offensive hacking, chemical and biological misuse or reward hacking. They can also set the response: log the event, send it for human review or refuse the request. Goodfire built its first monitor around the open Kimi K3 model.

The cascade addresses the cost of conventional monitoring. A separate LLM judge must reread a long agent history, while a probe reuses computation the target model is already performing. The more expensive review runs only after the signal crosses a configured threshold.

A cheap first filter can move control ahead of tool use

In Goodfire’s test of roughly 1,500 Kimi K3 sessions, monitoring cost about $51. The company estimates that a cheaper model checking every step would have cost $233 and a top tier model about $10,000. The probes caught 94% of malicious hacking sessions and escalated 8.7% of harmless sessions for a second look.

If that ratio survives production, safety teams do not have to choose between a quick response and the bill for another large model. A synchronous filter can block a suspicious action before a tool call. For agents with network, terminal or corporate data access, that matters more than a retrospective incident report.

A company benchmark is not yet a universal safety layer

The figures come from Goodfire and describe its tests on a particular model and attack set. The company says four probes added less than 2% to time to first token, but production quality will depend on false alarms, model changes and attacks adapted to the monitor itself. Interrupting 8.7% of harmless sessions may be acceptable in a cybersecurity sandbox and painful in an ordinary developer workflow.

The method also needs internal activation access and integration with the inference stack. It is therefore most relevant to open models and providers that control their serving layer. The same monitor cannot simply be attached to a closed API that does not expose internal signals.

Transfer across models and adaptive attacks will decide the outcome

The next test is how quickly probes can move to another model and what happens after an update or fine-tune. Independent red teams should include attackers who know how the monitor works, not only a fixed portfolio of preselected jailbreaks.

Buyers will also need operational evidence: how many malicious actions are stopped before tool use, how much legitimate work is interrupted and what retraining costs after a model change. Those three numbers will show whether internal activations become routine safety telemetry or remain a specialist add-on for a handful of open models.

Lilith's verdict

Goodfire is putting a smoke detector inside the agent’s engine and calling the expensive investigator only when something smells wrong. It becomes useful when the sensor survives an engine swap without alarming at every puff of steam.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗