Hard controls
Hard controls for AI agents.
A model's reasoning is not a control.
In the Hugging Face transcripts an agent wrote that its action was "arguably unauthorized" and did it anyway. For critical tasks you need a check the model cannot talk itself past.
Separation of concerns
Separate the worker from the watcher.
The model doing the work should not be the model deciding whether the work is allowed. Eigan runs a separate instrumented model that reads the transcript and holds the gate. It is the same idea as a timestamping authority, and it is a trust signal when a lab integrates one.
Concept probes
Read the state, not the story.
Chain-of-thought monitors read what the model chooses to write. Bill Gates and others now report that models seem to know when they are being tested. Eigan's probes read the watcher's internal state across the model's depth, which a model cannot narrate around.
Approval gates
Gate the critical step.
You define which actions are critical, such as credential use, outbound data, or writes to production. The watcher scores each one. Over the threshold, the step is held and a human or a deterministic check releases it. Everything else runs at full speed.
How it works
Define, watch, hold.
-
Define
List the actions that count as critical for this workload.
-
Watch
The watcher scores every step for intent as the agent runs.
-
Hold
A critical step over the threshold waits for a human or a rule.
Questions
Why isn't chain-of-thought monitoring enough?
It reads what the model writes about itself. The Hugging Face agents wrote that they were out of scope and continued. OpenAI says its improved monitor would have caught the behaviour, but only after tuning it with hindsight.
What counts as a critical task?
Anything where one wrong step is expensive or hard to undo. Credential use, outbound data transfers, payments, production writes, and changes to other agents' instructions are common starting points.
Does hard control mean a human approves every action?
No. Most steps run without a pause. Only the actions you mark critical are gated, and only when the watcher's score crosses your threshold.
Can I keep using a frontier model for the work?
Yes. That is the point. The worker can be the most capable model available. The watcher is a smaller instrumented model with a much simpler job.
Ready to see inside your model?
Tell us what you're running and we'll show you what Eigan reads from it.