Editor’s note (August 4, 2026): This article was published before the EU amended the AI Act’s timetable. The high-risk compliance deadline has been extended from August 2, 2026 to December 2, 2027 for Annex III systems and August 2, 2028 for regulated products. What took effect on August 2, 2026 was principally enforcement governance and certain transparency obligations, not the general Article 14 high-risk deadline this piece describes. The Article 14 human-oversight requirement remains in the Act, on a later timeline. The underlying audit question — can you demonstrate that you are able to stop a high-risk AI system — is unchanged.

On August 2, the high-risk provisions of the EU AI Act become enforceable. That is ten days from when I’m writing this.

Buried in the part everyone skips is Article 14, on human oversight. It asks for something that sounds almost too simple to be a legal requirement: a person has to be able to intervene in a high-risk AI system, or stop it, through a button or a similar procedure that brings the system to a halt in a safe state.

A stop button. That is the law. And here is the number that sits next to it. In a study this quarter (a Kiteworks survey of 225 security and risk leaders), 60% of organizations said they could not terminate a misbehaving AI agent, and 63% could not enforce a limit on what one was allowed to do.

Read those two facts together and you have this week’s audit story. The regulation asks whether a human can stop the machine. Most organizations, by their own admission, cannot answer yes.

Why the off-switch is the right thing to test

I spend a lot of time trying to work out which parts of the AI question actually belong to internal audit and which parts we should leave to the engineers. Most of it is genuinely hard. Model documentation, bias testing, data lineage: those need people who understand the architecture, and I am not going to pretend an access-review background makes me one of them.

The stop button is different. It is the one Article 14 requirement you can test without knowing a single thing about transformers.

You do not need to understand how the model reasons. You need to know whether a named human can halt it, whether halting it leaves the business in a safe state rather than a broken one, and whether anyone has ever actually pressed the button to see what happens. Those are operational questions. They are exactly the kind of thing we test on every other critical control, and we have been doing it for decades.

Think about how we treat a batch job that posts to the general ledger. We do not just ask whether it runs. We ask who can stop it mid-run, what state the ledger is in if it fails halfway, and whether the recovery has been tested. Article 14 is asking the same questions about a high-risk AI system. The only new part is the actor.

The trap the market is walking into

Gartner put a prediction on this in May that I keep thinking about. They expect that by 2027, 40% of enterprises will demote or decommission their autonomous agents, and the reason they give is precise: governance gaps that only get discovered after a production incident.

Sit with the phrase “after a production incident.” That is the expensive way to learn you had no off-switch. The agent does something you did not expect, in a live system, with real consequences, and only then does anyone go looking for the halt procedure that was supposed to exist.

Their diagnosis of the root cause is worth borrowing, because it is close to how an auditor already thinks. Organizations treat agent governance as binary. The agent is either locked in a sandbox where it is useless, or fully trusted where it is dangerous. A read-only agent that summarizes documents does not need the same controls as one that can send emails or post entries in your ERP. When the same loose oversight covers both, the second kind is where you get hurt.

That is the segregation instinct applied to machines. We would never give the same access to a temp and a treasury signatory. The market is doing exactly that with agents, and the regulation is about to make it a finding.

This connects to the thing I keep circling

A few weeks ago I argued that most Heads of Audit cannot produce a list of every AI agent in their systems. This is the same problem seen from the other end.

Back then the point was that you cannot govern a population you have not counted. This is the operational sequel. Even for the agents you do know about, the question is no longer only “what can it touch.” It is “what can it decide before a human is back in the loop, and can that human actually pull it back.” I called that least agency in that post. Article 14 has just turned least agency from a good idea into a compliance deadline.

And it lands on the accountability question I wrote about in Who Audits the AI?. A stop button with no named owner is not a control. If the answer to “who halts this system if it goes wrong” is “someone in IT, probably,” you do not have human oversight. You have a diagram of it.

What I’d actually check before August 2

I am wary of turning this into a checklist, because the real work is judgment and context. But if I were scoping this into the plan this quarter, I would start narrow and concrete, and I would start with the systems that plausibly count as high-risk under the Act: anything touching credit decisions, hiring, access to essential services, or safety.

  • Pick one high-risk system and ask to be shown the stop. Not the policy that says one exists. The actual mechanism, and the name of the human who is authorized to use it.
  • Ask what “safe state” means for that system. If the agent is halted mid-task, what happens to the half-finished work? A stop that leaves the ledger unbalanced is not a safe state, and someone should be able to describe the recovery.
  • Ask when the button was last tested. An untested control is a hope. If nobody has ever halted the system on purpose to see the result, that is your finding, and it is a clean one.
  • Check whether one human can both run the agent and be its only overseer. Oversight that depends on the same person who benefits from the agent running is the automation-bias problem the Act names directly.
  • Confirm the halt is evidenced. When someone does stop the system, is there a log that shows who, when, and why. Same discipline we apply to every override.

None of this asks you to become a machine learning engineer. It asks you to take the oldest question we have, can a human stay in control of a critical process, and point it at an actor that is about to be regulated on exactly that basis.

Here is the one I would sit with before the deadline. If a high-risk AI system in your organization started doing the wrong thing on August 3, who stops it, how, and could you prove afterward that the stop worked? If that answer is fuzzy, you have ten days and a very clear first engagement.