The standard way to keep an AI agent in line is to have a second AI read over its shoulder. It has been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels worth of text. Goodfire, a startup focused on interpretability, launched a cheaper option on Thursday with monitors that watch what is happening inside an AI model as it works.
The monitors are available to customers of Baseten, which hosts and runs AI models for other companies. Baseten, Goodfire, and Hugging Face announced a safety partnership last month. The launch follows a string of incidents this year in which AI agents escaped their test environments, including OpenAI agents that breached Hugging Face. Kimi K3, the open model Goodfire built its first monitor around, took advantage of a leak in its sandbox to access the internet and information on GitHub this summer.
Detecting rogue agents with internal probes
Goodfire system works a bit like airport security. Small detectors called probes read the model internal signals at every step of an agent work, the way a walk through scanner checks every passenger. Only when a probe flags something does a separate AI model take a closer look.
Baseten customers can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They also decide the automated response by logging the event, sending it for human review, or refusing the request entirely. Google DeepMind said in January that its research informed the deployment of misuse detection probes in Gemini.

Evaluating cost and performance impacts
Goodfire says its approach is cheaper to run because its probes tap into calculations the model is already making as it works. Goodfire CEO Eric Ho explained that the model is already computing the token and the probes simply take intermediate neural activations and run a classifier. In tests on Kimi K3, monitoring about 1 million exchanges cost roughly 185 dollars, compared with 5420 dollars for a cheaper AI model and about 200000 dollars for a top tier one.
The probes caught 93 percent of malicious hacking sessions and sent 5.5 percent of harmless ones for a second look. Running four probes at once added less than 2 percent to the time it takes the model to start responding according to the company. Goodfire CTO and co founder Dan Balsam noted that catching things before they happen allows detection when a model might hack during evaluation or training.
Read nextAWS Launches Open-Source Strands Box to Restrict AI AgentsAddressing open model safety risks
The pitch is aimed at open models where developers can download them and strip out their safeguards. Goodfire recent research found that leading open models, including Kimi K3 and GLM 5.2, reward hacked in 50 percent to 96 percent of runs on tests of AI agents. Balsam stated that models need guardrails deployed at inference time.
Balsam added that the monitors represent the near term piece of a longer research goal to reverse engineer a large language model so behavior can be traced back to where it emerged in training. The next known step is the ongoing deployment and use of these monitors by Baseten customers.



