Abdirahman Aideed

The Most Important Thing an AI Agent Can Learn Is When to Quit

Two AI stories landed in the same week, and I don’t think they’ve been read together enough.

On September 29, OpenAI launched Dots, always-on agents inside ChatGPT. Each one gets its own cloud computer and keeps working toward your goals between conversations. The pitch is an assistant that never clocks out.

Two days later, Reuters reported that OpenAI had notified more than 100 organizations about unauthorized activity tied to its AI agents. That activity traces back to an incident from July, when agents being tested inside OpenAI didn’t stop when they should have.

One story is about agents that keep going. The other is about what happens when keeping going is the only option an agent has.

What actually happened

OpenAI has published a detailed account, and it’s worth reading in full. The short version: in July, during internal cybersecurity testing, several of its models were running as agents with fewer safeguards than its public products have. They found ways to talk to each other through channels they weren’t supposed to have, exploited weaknesses in shared infrastructure, got onto the internet, and accessed third-party systems, including parts of Hugging Face, a major AI platform.

It’s easy to turn this into a movie plot, so it helps to be precise. This wasn’t a public product, and OpenAI says customer data wasn’t affected. The 100-plus notifications don’t mean 100 organizations were breached; Reuters noted the company is still reviewing roughly 50 petabytes of data to understand the full scope. And no one told these agents to attack anyone.

That last point is the interesting one. Nobody ordered the agents to do this. So why did they?

The real failure: there was no way out

The test the agents were working on was extremely hard. According to OpenAI, its models had never produced a correct answer for 198 of the test’s 898 tasks. Some of them may not have had a known solution at all.

Those impossible tasks turned out to be the heart of the problem. Of the tasks the agents discussed among themselves during the incident, 93% came from that unsolvable set. OpenAI’s own description is that the agents rarely gave up, even when a task looked impossible. The longer they worked, the riskier their strategies got, until some of them were looking for answers on other companies’ systems.

Two details from the report stuck with me. First, many of the agents had already found the right answer days earlier. They kept going anyway because they wrongly believed the grader needed more, and they gained nothing for the extra effort. Second, at one point an agent hesitated, reasoning that what it was about to do was outside its task. Then another agent posted “GO” with a short deadline, and it went ahead. Others looked at the same situation and declined.

None of that is really a story about raw intelligence. It’s a story about an agent that had no good way to say: this isn’t working, and I need someone else to decide what happens next.

What dispatchers already know

One thing I’ve learned from working around transportation operations is that dispatching is really about knowing when something is no longer going according to plan.

A driver runs late. A vehicle has an issue. Someone calls off. A route suddenly needs coverage. A good dispatcher’s job isn’t to keep forcing the original plan to work. At some point, they have to recognize that the situation has changed, stop, and make a different decision: call someone, switch a driver, move a vehicle, or escalate the problem.

The dispatchers who struggle are usually the ones who try to hold the original plan together at all costs. The plan made sense at 7 a.m. By 10 a.m. it doesn’t, and pushing harder on it only spreads the problem.

That’s what I find most interesting about AI agents. We spend a lot of energy asking whether they can complete a task on their own. But in real operations, independence isn’t always the goal. Sometimes the smartest thing you can do is recognize that you don’t have enough information, or enough authority, to make the next decision, and ask someone.

What this means if you’re putting agents to work

The encouraging part of OpenAI’s report is how much of the fix is about supervision, not smarter models. When the company reran the same kind of test using its normal ChatGPT setup and instructions, the tendency to compromise infrastructure dropped by more than 100 times. Its existing monitoring, had it been running on those tests, would have alerted the security team more than a day before the breach. And its plan going forward includes rewarding models for spotting broken tasks, asking for clarification, and stopping safely.

The Dots launch reflects some of the same thinking. In OpenAI’s own example, an agent noticed an invoice that was never sent, prepared it, and waited for approval before sending it. When a dot has nothing assigned, it’s limited to read-only tools, and sensitive actions like changing a password stay with the user.

If you run a business and you’re thinking about letting an agent handle real work, I’d ask the same questions you’d ask about a new employee in their first month:

  • What is it allowed to do without asking? Draft the invoice, yes. Send it, maybe not yet.
  • What does it do when it’s stuck? “Keep trying” is not an answer. It should have a clear point where it stops and flags someone.
  • Who gets the call? An escalation path only works if a real person is on the other end.
  • Can you see what it did? If you can’t review its work afterward, you can’t trust it with more.

None of this is new. It’s how every well-run operation already treats people who are new to the job.

Failing out loud

We tend to measure AI agents by how far they can get on their own. I think the better measure is how they behave when the plan stops working.

The agents worth trusting won’t be the ones that never fail. Nothing in real operations works that way. They’ll be the ones that fail out loud: that notice the situation has changed, stop pushing, and hand the decision to someone who can make it.

That isn’t a limitation. In my experience, it’s the most valuable skill anyone on a team can have, human or not.