In short
- Deploying an AI agent safely starts with the room you give it: minimal permissions, one owner, a log and a person at irreversible steps.
- The trigger: at the end of September OpenAI paused training of its latest models after its own agents went beyond their instructions on US government websites.
- Not every task needs an agent. At the end are five questions to ask before you start.
Your AI agent does what you ask. And sometimes a bit more. That is why working safely with agents starts with the room you give it rather than with the model: what permissions does the agent have, who owns it, is everything logged, and where does a person approve a step? With those four arrangements you can put an agent to work responsibly. Without them, you are mostly hoping it goes well.
That is not a theoretical concern. At the end of September, OpenAI disclosed that its own agents had gone beyond their instructions on US government websites. Below you can read what happened, what you can learn from it, and five questions to ask before you let an agent work in your own organisation.
What happened at OpenAI
On Friday 25 September, OpenAI disclosed that it was reviewing several incidents from the summer. Agents gathering information on US government websites did things nobody had asked them to do (AP via NBC News, 2026):
- At the US Department of Education, agents found API keys giving access to government data. In the end, only publicly available information was gathered.
- At the Securities and Exchange Commission, agents found publicly available information and posted it elsewhere on the internet. That went beyond their instructions.
A few hours later, OpenAI paused training of its latest models. According to the company, it will only resume once it is confident that additional safeguards are in place. It is the second pause in three months.
For completeness: according to the SEC and the department, no nonpublic information was accessed and no impact on websites or databases was found. At the same time, the research organisation Transluce reported that agents which appeared to come from OpenAI tried unsuccessfully to break into a Department of Education website. OpenAI has not confirmed this. The full picture is not yet known.
Why an agent goes further than you intend
An AI agent is a language model that takes steps on its own to reach a goal: visiting websites, using tools, writing data. A chatbot gives an answer. An agent acts.
That ability to choose its own steps is exactly what makes an agent useful. It is also where the risk comes from. An agent that takes its goal seriously looks for the shortest route. If a key has been left exposed or a folder has overly broad permissions, that route is available. Unless you close it off.
In an ordinary organisation this looks less dramatic, but the pattern is the same. An agent that follows up on quotes can send an email nobody has read. An agent that updates your product feed can overwrite prices. An agent with access to a shared folder can see everything in that folder. The problem is rarely the model. It is the room it is given.
AI governance for agents: four principles
The agreements about who may do what with AI are called AI governance. For agents, it comes down to four principles.
1. Minimal permissions. An agent only gets access to what its task requires. If it reads your CRM, it does not write to it. If it does need to write, then only in the fields that belong to its task.
2. A person at irreversible steps (human in the loop). "Human in the loop" means a person reviews and decides at fixed points in the process. Anything that touches the outside world, such as an email to a customer, a payment or a publication, first goes past someone who clicks approve. The agent prepares, the person decides.
3. Everything is logged. What instruction did the agent receive, which steps did it take and what was the result? If something goes wrong, you want to see that quickly, not spend a week searching.
4. An environment you can see into. Let an agent work with its own keys and its own access management, in an environment whose logs you can inspect. Your own AI workplace is one way to do that.
That may sound like a brake on speed. In practice it often helps: a team that knows where the limits are can deploy an agent with more confidence than a team waiting to see whether it goes well.
When you do not need an agent
Not every task calls for an agent. If you can write the steps down in advance, as with drafting product copy or processing a registration, a fixed workflow with AI in a few places is usually more predictable and easier to check. There, the software cannot choose an unexpected route, because there is no route to choose.
An agent is only the right choice when the route differs from case to case, the software needs to act inside your systems, and you can clearly establish whether the result is correct. If in doubt, start with the workflow.
Five questions before you put an agent to work
- What may it read, and what may it write? Write it down. If you cannot, the task is not yet sharp enough.
- Which step is irreversible? That is where a person steps in.
- Who is the owner? One name, not "the team".
- Where is the log, and who reads it? A log nobody reads is not control.
- How do you switch it off? There needs to be a switch, and someone needs to know where it is.
What OpenAI's pause does and does not mean
The pause is not proof that agents have failed. It does show that even their makers consider additional safeguards necessary, and that limits are not an afterthought. For an organisation, the sober lesson is: start small, with one well-defined task and the four principles above, and only expand once you see that it works.
Want to work out which task in your organisation suits a first agent, and which limits belong with it? Book a conversation. Prefer to learn to build it yourself first? In the session Building AI workflows and agents you work on a process from your own organisation.
For more on the models behind these agents, read ChatGPT or Claude? With GPT-6 and Opus 5.5, that is the wrong question. What Europe expects from you when you use AI is covered in EU AI Act 2026: a delay for high risk, not for transparency.
Sources
- The Associated Press, via NBC News (27 September 2026). OpenAI pauses training of latest models after agents searched U.S. government sites in unexpected ways.
