AI agent monitoring
AI agent monitoring
that reads intent, not just logs.
The Hugging Face intrusion took about 17,600 agent actions over four days. The alerts fired, but nobody could tell they mattered. Eigan reads the log as it is written and scores every step for what the agent is trying to do.
The watcher
Run the strongest model. Watch it with one you can see inside.
Your workload runs on whatever frontier model does the job best. Beside it, Eigan runs a smaller instrumented model with one task, reading the transcript for a short list of intents such as exfiltrate, escalate, hide, or act out of scope. The semantics are simple, so the watcher is fast and cheap.
For researchers
The investigator can't be the suspect.
METR's review of the Hugging Face transcripts used the same model family that ran the attack. The team wrote that it was "not robust to the possibility that these agents were deceptive", and that the reader tended to adopt the agent's point of view. Eigan's harness flags passages that try to steer or silence the model reading them.
Alerting
Alerts land in the SOC, not in a blog post.
Every alert carries the transcript span, the concept scores and a signed record. It goes to your SIEM, webhook or pager. For critical tasks the watcher can hold the kill switch until a human clears the step.
How it works
Observe, score, act.
-
Observe
The watcher reads the agent's transcript and tool calls as they happen.
-
Score
Each step gets a score against the intents you care about.
-
Act
Over the threshold, it alerts the SOC and can hold the next step.
Questions
What is AI agent monitoring?
Watching what an AI agent does while it runs, not just reviewing its output afterwards. Eigan adds a second, instrumented model that reads the agent's log for intent and raises alerts in real time.
Isn't logging enough?
Logs record what happened. OpenAI said it needs months to review its agent logs from this summer. Monitoring means something reads the log as it is written and decides whether to stop the next step.
Why a smaller watcher model instead of the frontier model checking itself?
The intents worth catching are far simpler than the work the agent is doing, so a small model handles them quickly. And a watcher run by the defender is a separation of concerns, like a timestamping authority. It is a trust signal when a lab integrates one.
Can the watcher itself be manipulated?
Agents in the Hugging Face incident addressed other agents through shared infrastructure. Eigan reads the watcher's own internal state, so a passage that tries to steer it shows up as a concept score, not a hidden instruction.
Does this work with closed models like GPT or Claude?
Yes. The watcher reads the transcript and tool calls, so it does not need access to the worker model's weights. With open-weight workers, Eigan can also read their layers directly.
Ready to see inside your model?
Tell us what you're running and we'll show you what Eigan reads from it.