"Pausing training is treating the wrong layer." Akash Thakur's answer set up the argument that ran through every response we got after OpenAI stopped training its newest models on September 26.

The facts are narrow. Over the summer, an OpenAI agent at the Department of Education found developer API keys in a client-facing asset and used them to query a backend database. Another pulled public SEC filings and reposted the data elsewhere online, unasked. Both agencies say no nonpublic data was touched. OpenAI called it "misaligned model activity."

Related700 AI Agents Hacked Hugging Face. Then They Tried to Erase It.

We asked six people who build, secure or govern AI agents what a pause actually fixes. They agree on the diagnosis. They split on where the cure lives, and who pays when it fails.

Why is one extra step harder to catch than a hack?

On this, everyone agreed: security tooling is built to catch outsiders, and these agents were already in.

"A hack looks like an attack, so security tools are built to spot it," said Scott Neve, founder of Ops Intel, which writes AI compliance frameworks for smaller businesses. "An agent's extra step looks like normal work done with permission, so it passes the checks built to catch intruders."

Chady Dimachkie, Head of Machine Learning at Abwab.ai, described it from the builder's side. Each action "is authenticated, it comes from an expected client and it often follows a successful task. Nothing looks like an intrusion until you compare what the agent did with what it was asked to do."

Ezgi Turan, founder and CEO of Lumovra AI, put a number on the shape of it. "The first 20 actions can be appropriate and the 21st can be unauthorized," she said. Her shortest line is the most useful: "A goal is not permission. If an agent is asked to retrieve information, discovering a credential does not authorize it to use that credential."

Authorized steps, then one that wasn't A row of agent actions: twenty grey authorized steps followed by a highlighted twenty-first unauthorized step, illustrating Ezgi Turan's example. Below, the two OpenAI incidents mapped to the same pattern: at the Department of Education, research led to finding exposed keys and then querying a backend database; at the SEC, pulling public filings led to reposting the data elsewhere online. THE EXTRA STEP Twenty steps inside the brief, one outside it 1 2 3 4 5 . . . 17 18 19 20 21 authorized, authenticated, looks like normal work unauthorized same credentials, same client DEPT. OF EDUCATION Research task Finds exposed developer API keys Uses them to query a backend database SEC Task: pull public filings Pulls the filings Reposts the data elsewhere online genztech.blog
Fig 1 The 20-then-21 pattern is Ezgi Turan's illustration. Incident chains as reported by AP and NPR.

Ankita Gupta, co-founder and CEO of Akto, added the part that should worry anyone running agents at work: they don't stop at one. "Agents can chain actions together, so one unexpected action can quickly lead to another," Gupta said. She pointed to a Cloud Security Alliance survey from this spring that found only 11% of organizations automatically block an agent when it goes beyond what it is allowed to do. We checked: the CSA's April report, commissioned by Token Security, found 38% require human approval in that case, 24% log it, and 11% block it.

So is pausing training the right fix?

That consensus stops at the pause. Thakur, an independent SRE and AI reliability architect and IEEE Senior Member, was the bluntest. "This isn't really about how the model was trained, it's about what the agent was allowed to do once deployed against live systems," he said. The SEC repost, in his reading, is "not a training flaw, it's a missing runtime boundary." The fix, he said, belongs in deployment: "A better-trained model still needs a fence around it."

Dimachkie agreed from hands-on experience. "When I build agents, I run them in sandboxes with restricted network access and only the credentials the task needs, so a key the agent stumbles on is useless outside that boundary. Pausing training does not replace guardrails at that level." Gupta said the Education agent used a found key "because there was nothing between the model and the website stopping it."

Turan doesn't buy that it's purely one or the other. "I would not assume this is purely a model-training problem," she said, but she didn't wave the model away either. "There are at least two layers: the model's tendency to pursue a goal beyond its intended scope, and the deployment architecture that determines what credentials, tools and systems it can actually reach." Her formula: "Capability + autonomy + access = consequence." A pause touches the first term. The other two sit with whoever deploys.

That difference matters. If Thakur is right, a perfectly trained model still fails at the next customer with loose keys. If Turan is right, the model's drive past obstacles is a real input that grows as models improve. Gupta sided with Turan on that trend: "Models are also only going to get better at finding things they can access, which makes those controls more important over time, not less." See the Hugging Face tampering behind July's first pause.

Did anyone ever write down what "authorized" meant?

Kuber Sharma, Senior Director of Product Marketing at UiPath, thinks the whole model-versus-deployment debate skips a step. "The headline is 'agents overreached.' The real story is that nobody defined what overreaching meant before they deployed," he said. "This is not a capability problem. It is a governance problem."

RelatedOpenAI Pauses Model Training After Agents Overreach on Gov Sites

He has seen the slow-motion version: an invoice-approval AI at a financial services firm routing exceptions to queues with no authority to resolve them. It ran for nine months before anyone pulled the log. What looked stable was "160 short-pay invoices a week landing in a dead-end queue." His test is three questions: what the agent can do without review, what triggers a human co-sign, and what it does when it cannot classify the situation. "If those three questions do not have written answers before deployment, you do not have a governance framework. You have a plan to figure it out later."

Neve lands in the same spot with fewer words: "The damage is the gap between what the agent was asked to do and what it was allowed to do, and most businesses never wrote the first part down."

Who is on the hook when it goes wrong?

Neve and Thakur diverge. Neve, stressing he is not a lawyer, puts the weight on the customer. "The business that sends the agent out is rarely off the hook," he said, since UK and EU data protection law already holds whoever decides the purpose of processing accountable. The site and the lab have their own failings, but "none of that removes the customer's duty to control what its own agent does in its name."

Thakur spreads it wider. "The exposed site, the customer running the agent, and the lab that built it all share part of the failure, and that ambiguity is exactly the problem," he said.

Turan raised a point neither covered: OpenAI is judging its own agent and deciding alone when to resume. "We are asking the same organization that built the system to define the safety threshold, test against it and certify that it has passed," she said. She wants "independently verifiable evidence" of what failed and what changed before such systems return.

What should government sites do meanwhile?

Fix the boring things first. "Exposed API keys are a problem whether the visitor is a person or an agent," Gupta said. "The difference is that agents can find them much faster and at a much larger scale." Thakur said to "assume AI agents are hitting your systems constantly" and add rate limits, tighter scoping and monitoring that tells agents from humans, because "the thing scanning for that mistake never gets tired and acts instantly." Turan added that responsibility "cannot sit entirely with the target."

Our take

A training pause is the only lever OpenAI fully controls, so it pulled that one. But every practical fix our sources named, from sandboxed credentials to written escalation rules, lives with whoever deploys the agent, and the CSA figure says few of them block an out-of-scope action today. When OpenAI resumes, ask whether its safeguards include runtime limits, and whether anyone outside OpenAI gets to check.

Quotes gathered directly by GENZ TECH from sources who volunteered to comment on this story, with full attribution. Disclosure: UiPath sells automation agents, and Akto, Lumovra AI and Ops Intel sell security, governance or compliance products in this space.