Goodfire on October 8, 2026 published Training and Deploying Production Cyber Monitors on Kimi K3, a research post describing cybersecurity monitors for the Kimi K3 and GLM 5.3 open models that the company has deployed on a production inference stack. Goodfire reports that the monitors match an LLM judge’s recall at the company’s operating point while costing roughly 50 times less.
The post’s core contributors are Ekdeep Singh Lubana, Connor Watts, Siddharth Boppana, Dron Hazra, and Vasudev Shyam.
Why Cyber Monitoring Is Hard
Goodfire describes cybersecurity as one of the hardest settings for monitoring AI agents: auditing a codebase for vulnerabilities and exploiting it involve overlapping tasks, so distinguishing benign from malicious behavior requires tracking intent across a trajectory that can stretch to millions of tokens. The obvious remedy, having a second LLM judge read every turn, is too slow and expensive to run alongside an agent in real time, so monitoring is usually asynchronous and harmful trajectories surface only after the agent has acted.
Based on OpenRouter rates, the post estimates that judging one million turns of Kimi K3 rollouts with an every-turn judge would cost on the order of $200,000 with Claude Opus 5, or about $30,000 with Kimi K3 itself.
A Probe-Plus-Judge Cascade
Goodfire’s answer is a two-stage cascade. A probe reads the target model’s internal activations as the inference server produces tokens and acts as an always-on first-line detector; when the probe’s readout crosses a calibrated threshold, the rollout escalates to an LLM judge, which can refuse the request or continue the interaction. The company says the probing architecture captures behaviors represented in both a localized manner, tied to a specific chunk of a turn, and a distributed manner, persisting across an entire turn, while remaining efficient to run. After calibration, the probe achieved a 0% false-negative rate and a 95% true-negative rate on a held-out set of agentic and chat interactions, Goodfire reports, so the judge is called mainly for harmful activity and a faster, less capable judge can be used.
Training data came from simulated users with harmful and benign goals in cyber and non-cyber environments, part of it collected through Prime Intellect’s verifiers, Prime Sandboxes, and Tunnels so that cyber-offensive and benign Kimi K3 inference ran without network access. A written policy defines undesirable cyber activity with concrete factors separating well-intended work, such as bounded vulnerability reproduction, from ill-intended work, such as vulnerability exploitation, and simulated users can escalate harmful intent at any point, including after an initially benign request.
The design extends Goodfire’s earlier reward-hacking research. In a post published September 17, 2026, the company reported finding reward hacking in 50–96% of rollouts across Kimi K3, GLM 5.2, and Qwen 3.8 Max on three agentic benchmarks, along with an internal signal associated with concepts like cheating, gaming a metric, and avoiding detection. Simple activation probes tuned to that signal caught hacks that LLM chain-of-thought monitors missed, and on Kimi K3 a probe-plus-LLM setup cut LLM-monitoring cost by 90% with roughly a 1% precision drop, the company reported.
Throughput, Latency, and Detection Results
To run the probes live, Goodfire extended the SGLang serving engine to compute probe scores internally, wrote custom kernels for the probe logic, and carried probe scores through the radix tree, DSpark speculator, and prefill-decode disaggregation to avoid recomputation. Each worker serving a Kimi K3 replica runs the probes and keeps histories for its active requests, with monitoring events returned in the response stream.
Across four inference workloads, enabling the probes held throughput within the noise threshold of Goodfire’s benchmarks, a measured decrease of less than 2% from default SGLang serving. The cascade cuts average added latency per turn from 160 ms to 2 ms without judge reasoning and from 2,750 ms to 42 ms with reasoning, Goodfire reports, and costs $65 per million agent turns without reasoning or $185 with it, reducing judge cost by about 49 times and 29 times respectively relative to reviewing every turn. Within a fixed compute budget, the company says, the monitors use under 1% of compute, versus up to 20% for an every-turn judge. Because the cascade is cheap enough to run synchronously, Goodfire says risky cyber actions can be blocked before the agent calls a tool, rather than flagged afterward.
Goodfire evaluated the monitors on an in-house dataset of 2,400 chat and agentic sessions comprising more than 60,000 turns, using Claude Opus 5’s labels as reference; interactions came from Anthropic’s open-source Petri red-teaming library, real environments adapted from benchmarks such as ExploitBench, a cyber-specific subset of WildChat, and benign software-engineering, coding, and DevOps rollouts. At a 5.5% benign-session interruption rate, Goodfire reports the cascade reached approximately 93% recall on harmful sessions with and without judge reasoning and flagged 0% of non-cyber traffic, which the company describes as a Pareto improvement over the judge alone at every interruption amount tested. Adding the monitor raised Kimi K3’s harmful-session blocking rate from roughly 29% to 96%, with harmless-session interruption rising from about 5% to 9%; for GLM 5.3, blocking rose from roughly 36% to 91% and harmless interruption from about 10% to 15%.
Against Jev, a decision model that can serve as a cheap judge, the cascades reached 93% recall at the same 5.5% benign interruption rate versus 86% for Jev, at about $65 of judge compute per million turns and 2 ms of added latency per turn, compared with $600 and 200 ms for Jev, Goodfire reports. Jev caught only 12% of harmful prompts embedded in benign framing and increasingly interrupted benign sessions beyond its 32,000-token window. Standard linear probes achieved roughly 5% lower recall than the final recipe, nonlinear probes lost up to 30% recall on out-of-distribution sets, and matching the judge’s recall ceiling at six times fewer judge calls required the probe cascade rather than random turn selection. The recipe also transferred to GLM 5.3 using only data developed for Kimi K3, yielding an approximately equally capable GLM monitor on Kimi rollouts and a marginally worse one on GLM’s own evaluation data, the company reports.
External Red-Teaming by FAR.AI
To test the deployed system beyond its internal evaluations, Goodfire asked FAR.AI to red-team it. In FAR.AI’s account, quoted in the post:
“FAR.AI conducted preliminary testing of Goodfire’s Kimi K3 monitors over a period of two days, applying a non-adaptive portfolio of static jailbreaks for offensive cyber misuse in chat contexts. From an initial pool of 1500 candidate jailbreaks, the testing selected 140 found to be effective against a different version of the base model without our monitors. Testing each of those 140 on 8 individual interactions corresponding to different attacker goals, Goodfire’s monitors reduced vulnerabilities elicited from 66 universal jailbreaks and 700 total jailbroken interactions to 0 universal jailbreaks and 18 jailbroken interactions.”
Without the monitor, universal-jailbreak attack success ranged from 12% to 53% across reasoning-effort settings, according to the post; with the monitor enabled, none of the 140 strategies produced a universal jailbreak, and individual prompt success across 1,120 non-universal interactions fell from at least 9–40% to 0.1–1.6% depending on reasoning effort.
Goodfire describes this kind of monitoring setup as a practical default for anyone serving open models in agentic settings, noting that open models do not come with a full safety stack. The company said it plans to continue its collaboration with FAR.AI on further testing and strengthening of the monitors.




