Padlock illustration next to the title

How to get an OpenClaw or Hermes agent approved by your security team

• 11 min read

I’ve found that when anyone responsible for security hears “OpenClaw” or “Hermes,” they almost have a heart attack. The conversation can start and end with “absolutely not.”

So, before mentioning either technology, come to security with the intention of building on agreements. Start here:

Supporting unattended agents in our business could give us a significant competitive advantage as a company. Let’s work together and figure out how we can make this work.

Give them a business problem worth solving. Then introduce the technology.

In my case, I said I wanted to use OpenClaw or Hermes, followed by:

I know there are a lot of security concerns here, and this is potentially dangerous technology. Let’s work together, come up with all of your concerns, and address each one so we can actually deploy this to production and get it working confidently.

Now comes the hard work.

This conversation will be a lot easier if you focus on one discrete, isolated use case with as little access as possible to PII or sensitive company information. “I want an agent that can do anything” is a pretty efficient way to get rejected.

This agent eventually became known as Nines in our organization. I’m using Hermes for Nines because, in our setup, it has been more stable with concurrent long runs and multiple active threads.

Bring a one-screen approval packet

Put the scope in writing before the meeting. Here’s a compact template using Nines’s workflow. Fill in the exact identities, destinations, limits, and retention periods for your deployment. “We’ll figure that out later” is not a particularly compelling security control.

Use case:

  1. When a new P1 or P2 customer bug report comes in, automatically triage that ticket.
  2. Try to reproduce it.
  3. Check the logs and metrics to validate the report.
  4. Post an analysis on the ticket.
  5. If possible, propose a fix as a GitHub PR.
ItemWhat to put in the packet
Data in / data outIn: Slack requests, Jira tickets, approved code, scoped logs/metrics. Out: Slack replies, ticket analyses, PRs. Name permitted data classes and model provider.
Identity and credentialsDedicated identities, scopes, credential lifetimes, revocation owner. No human SSO reuse.
Tool allowlistSlack read/write; approved repo clones and PRs; Jira read/update; logs/metrics read-only. List exact channel, project, and repo restrictions.
Network allowlistExact service/model destinations and ports; Squid policy; firewall blocks direct outbound access.
Isolation diagramVM contains agent → Squid → allowlisted destinations. Diagram below shows the boundaries.
Kill switch and time/spend limitsStop/revoke procedure and owner; runtime cap; per-run/daily spend caps. Enforcement outside the agent.
Log destinations and retentionDestination, captured events, redaction, readers, retention period. Nines uses our main logging system.
Supply chainNines: community skill installation disabled. Document approved sources, pinned digests, review/signature policy, and update owner for any skills you ship.
Test planUnauthorized trigger rejected; blocked destination unreachable; schedule change denied; credential revocation stops subsequent access. Attach evidence.

Nines had a specific task, specific inputs, and specific outputs. Now we had something to talk about.

The concerns my security team brought up were access control, prompt injection, deployment isolation, protection of PII, and rogue execution. I found these conversations much easier to navigate once we broke them down and talked through them individually.

Let’s take each one in turn.

1. Access control

Treat the agent like a service account with worse judgment.

The permissions conversation should still sound familiar. You don’t give an engineer access to the company bank account. Likewise, you don’t give finance admin access to GitHub.

We’re following the same principles here. Give the agent the least possible permissions and explicitly limit the tool operations it can perform.

Give it a separate identity, with short-lived credentials wherever the service supports them. For services that require longer-lived tokens, scope and rotate them, and test revocation. Don’t reuse your human SSO session or mount your home directory into the container. Your browser cookies and SSH keys are not part of the onboarding package.

When I proposed Nines to security, I listed the exact inputs, outputs, and tool calls it would have:

  • Read and write messages in one specific Slack channel.
  • Clone repositories from an explicit list of approved GitHub repositories.
  • Open pull requests.
  • Read Jira tickets and update them.
  • Access our logging and metrics systems with read-only permissions.

That list gave us something concrete to review together. Security could look at each capability, ask why it was needed, and discuss how to restrict it. “It needs access to GitHub” leaves a lot to the imagination. “It needs to clone these repositories and open PRs” is a much easier conversation.

Make sure those limits are enforced through credentials, service permissions, and tool configuration. Writing the list into the prompt is not access control.

Opening a PR makes sense for this workflow. Give humans responsibility for approving and merging it, and enforce that through repository permissions and branch protection.

You can also limit network access using something as simple as a Squid proxy. Pair it with firewall rules that block direct outbound access, so using the proxy is mandatory. Otherwise, you’ve mostly given the agent a suggestion.

2. Prompt injection

Start with the inputs. Who can chat with it? Where does the information come from?

In my case, the inputs were through a private chat channel accessible only to certain users, and Jira. Only tickets submitted by our own support engineers would be processed.

That reduces exposure. There is still a catch: a support engineer can paste customer content into a ticket, and logs can contain text supplied by customers. An internal ticket doesn’t magically make everything inside it trustworthy. OpenClaw’s own documentation calls out this risk.

This is where the access controls earn their keep. Treat ticket contents and logs as material to investigate, and enforce what the agent can do through its tools, credentials, and network access.

If a ticket tells the agent to upload company data to some random website, the network should block it. “Please ignore malicious instructions” is doing a lot of heavy lifting if that’s your entire defense.

3. Deployment isolation

Deploy it in isolated infrastructure, in its own project or account.

I deployed mine in a Docker container on a sandboxed VM, with Squid running in another Docker container on that same VM. The sandboxed container was one of security’s requirements. Here’s that layout with the boundaries you should enforce:

Both containers share the VM, so protect the host too. The agent must not be able to change the proxy policy or firewall rules.

Docker isn’t a magic security force field. How you configure it matters.

4. Protection of PII

The starting point is the same as access control: if it doesn’t need access to something to complete its task, don’t give it access.

For this workflow, pay attention to tickets and logs. Both can contain customer information. Limit what the agent retrieves, redact sensitive fields before they reach the model, and agree with security on which model provider can process the remaining data.

Also pay attention to the outputs. Copying sensitive information from a restricted log into a widely visible Jira comment is still a leak. The fact that the agent was being helpful does not improve the situation.

5. Rogue execution

This is probably the biggest one. These agents can take a request and get a little too creative about what “helpful” means.

OpenClaw’s heartbeat is one way it can run periodically without a new message from you. If your workflow doesn’t need that, disable the recurring heartbeat. OpenClaw documents heartbeat.every: "0m" for that purpose. Other event-driven execution can still occur, so review those triggers too. OpenClaw heartbeat documentation.

For a workflow like mine, you can put concrete limits around that behavior:

  • Control what starts a run. Accept only the approved ticket events and authorized chat requests. Disable unused schedules and integrations. Have the receiving service validate the sender, project, and event type before starting the agent.
  • Prevent it from expanding its own job. Remove unnecessary scheduling and configuration tools. Keep the configuration that controls its permissions and triggers outside its writable workspace, enforced by filesystem permissions. If it has shell access, account for that too—hiding a tool won’t help if it can make the same change from a terminal.
  • Restrict where its work can go. Enforce allowed ticket projects, repositories, and network destinations in the services it calls. Let it propose a fix. Require human review before that fix can be merged or deployed.
  • Put a hard stop on the work. Use a supervisor outside the agent to enforce a deadline and terminate the run and its child processes when time is up. Apply resource and spending limits outside its control, and keep a way to stop the workload and revoke its credentials.

Those controls give security something they can inspect and test. Trigger an unauthorized request. Try a blocked destination. Attempt to change the schedule. Verify that the boundary holds.

The agent has a bug ticket to investigate. The agent doesn’t need a side hustle.

6. Supply chain

Expect security to ask about the skill marketplace. A community skill can bring instructions, scripts, and dependencies into the agent’s environment. OpenClaw’s documentation says to treat third-party skills as untrusted code.

For Nines, the policy is simple: community skill installation is disabled. It has a defined job. It doesn’t need to go shopping for new capabilities halfway through a bug investigation.

For a narrow workflow, I’d recommend the same approach: ship only the skills the task needs. Enforce that through deployment controls: keep approved skills read-only to the agent, prevent it from loading skills from writable locations, and block runtime installation paths. Removing an install button is not enough when the agent has a shell.

If you need community skills, review their instructions, scripts, and dependencies before deployment. Pin the reviewed content to immutable versions or digests. Have your build pipeline package that content into an internally signed artifact, and make deployment reject unsigned artifacts or signatures from unapproved signers. Keep signing keys and deployment permissions out of the agent’s reach. Review updates through the same process.

A signature tells you where the package came from. Someone still has to read what it does.

Security advisories land regularly, and not all of them get a CVE. Watching CVE feeds alone isn’t enough.

You don’t need a new process for this. We brought the agent into the ones we already run:

  • Vulnerability management. We pulled the agent’s repository, container images, and dependencies into our existing scanning. The version is pinned, and upgrades get reviewed like any other PR. Keep an eye on upstream security advisories too, including the ones without a CVE.
  • Incident response. AI-related incidents are covered in our incident response runbooks, so we know what to do if something goes wrong. An unattended agent is a good reason to have that worked out before you need it.

Auditability and the compromise

Security also required auditability, and this was a meaningful part of getting through the conversation.

All of our agent logs, prompts, and outputs go into our main logging system. If something unintended happens, we can go back and audit what it actually did: what it was asked, what information it received, and what it produced.

For anyone setting this up, include tool calls and their results in that trail, and protect the records from alteration by the agent. Apply redaction and access controls to the logging pipeline too. You still don’t want credentials or unnecessary PII in your logs.

Auditability won’t prevent every mistake. It does give you a way to investigate one without asking the agent to write its own incident report from memory.

That was part of the compromise: I wanted the agent running unattended; security required sandboxing and visibility into what it was doing. Those requirements became part of the deployment.

Bring the packet. Name the owner. Test the boundaries. Show security what happens when a trigger is unauthorized, a destination is blocked, a schedule change is attempted, or a credential is revoked.

That gives security a concrete decision to make: whether this agent, doing this job, with these permissions and these controls, is acceptable to run.

Related resources

European data residency for calendar APIs: how to evaluate vendors

Calendar API evaluations in Europe can go wrong at the first step, and it is…

What it actually takes to migrate a scheduling integration

Vendors describe migrations as straightforward. Engineering teams who have run one describe them differently. The…

Data residency considerations for scheduling and calendar features in Europe

When a scheduling vendor consolidates its regional environments, every customer built on that region inherits…