Ask your CIO what determines whether an AI agent is safe to deploy in your hospital. Most will describe an accuracy score. A vendor bake off. A pilot with encouraging results. Almost none will describe the system that catches the agent when it is wrong.
Changi Airport Group runs one of the busiest airports on earth, and its technology leader admitted something this week, in a room full of health IT executives, that should have made every one of them uncomfortable. Speaking at HIMSS26 APAC, a CAG consultant said his team has walked away from multiple autonomous AI projects because the organization could not prove they were safe to run unsupervised. Not because the model underperformed. Not because the use case lacked value. Because nobody could guarantee the system would stay inside acceptable limits once a human stopped watching it.
Sit with that for a moment. An organization with the discipline to admit a project is not safe, and the discipline to kill it anyway, after the demo already worked. Most healthcare organizations cannot do that. Here is why, and it has nothing to do with technology.
The demo is lying to you, and your brain wants to believe it
A polished demo triggers a specific mental shortcut. Psychologists call it the fluency heuristic. Anything that feels smooth and effortless to process gets read by the brain as true, safe, and trustworthy, even when smoothness has nothing to do with actual safety. A vendor spends months polishing a demo to feel smooth. Your instinct reads that smoothness as evidence. It is not evidence. It is a magic trick, and executives who would never fall for a magic trick on stage fall for this one in a boardroom, every week, because the trick is aimed at exactly the part of the brain that skips scrutiny when something feels easy.
Ask yourself when you last approved a technology because the demo felt impressive rather than because someone showed you the failure modes. If you cannot remember, that is the answer.
Normalization of deviance, and your organization is doing it right now
Sociologist Diane Vaughan coined the term normalization of deviance studying the Challenger disaster. The idea is simple and it applies far beyond aerospace. An organization skips a safety step once. Nothing bad happens. So it skips the same step again. Each skip without consequence quietly redefines what counts as acceptable, until the organization is running well outside its own safety margin and nobody notices, because nobody remembers where the line used to be.
Healthcare is doing this with agentic AI right now, in plain sight. A team signs a contract for an agent that routes prior authorizations. Identity governance gets pushed to phase two because phase one needs to ship. Nothing breaks in week one. Phase two quietly disappears from the roadmap. Six months later the agent is acting inside financial and clinical workflows with no verifiable identity, no access boundary, and no audit trail, and nobody in the building remembers deciding that was acceptable. Nobody decided it. It simply happened, one skipped step at a time.
The central healthcare AI problem is not model selection. It is whether an organization has the workflow controls, data foundations, escalation paths, monitoring, identity controls, and human accountability required to safely delegate a real task.
Why nobody owns the outcome
Put a vendor, an IT director, and a clinical leader in the same room to approve an agent, and you would expect three people watching closely. What you actually get is closer to nobody watching at all. Social psychologists call this diffusion of responsibility. When several people are technically in a position to act, each person quietly assumes someone else has it covered, and the more people in the room, the less any single person feels personally on the hook. Apply that to an AI steering committee and you get exactly the failure mode healthcare keeps producing. Everyone signed off. Nobody owns what happens next.
Changi Airport Group does not have this problem, because more than sixty percent of its AI experiments never reach deployment, and that rejection rate is treated as the system succeeding, not failing. Compare that to a hospital that already announced its AI initiative at a board meeting and already spent the budget. Killing the project now means admitting the announcement was premature. Robert Cialdini described this decades ago as commitment and consistency bias. Once you have told the room you are doing something, you will defend the decision long after the evidence says you should not.
Who on this team is responsible for catching the agent when it is confidently wrong, and what do they actually check before the output reaches a patient, a claim, or a supply chain? If nobody can answer that with a name and a process, the organization has adopted a tool without adopting the discipline that makes it safe.
The six question test
Do not take my word for any of this. Run your own organization through six questions, right now, before you read any further.
Where exactly does the agent's authority end? If the answer lives in a sentence in a policy document rather than a boundary enforced in the system itself, the agent has no real lane. It has a suggestion nobody is required to follow.
What happens when the agent is fed bad data? Most health systems still run on fragmented records, disconnected department systems, and manual reconciliation. An agent built on top of that fragmentation does not fix it. It reaches the wrong conclusion faster and with more confidence than a person would.
Who does the agent call when it does not know? Every agent needs a documented moment where it stops and hands control back to a person. If that moment only got defined after the first incident, you already know how this story usually ends.
Can you see what the agent decided this week, right now, without calling the vendor? An agent that runs without continuous visibility into its own decisions is not autonomous. It is unsupervised. Those are different words for the same liability.
Does the agent have its own identity? A verifiable identity, an access boundary, and an audit trail, the same way a new hire gets on day one. Most healthcare IT stacks were never built to issue identity to a piece of software making independent decisions, and that gap is exactly why it keeps getting skipped.
Whose name is on the incident report? Not a vendor's terms of service. A named person inside your organization who can explain, defend, and if necessary reverse what the agent did. Airports figured this out because a mistake grounds a plane. Hospitals need the same clarity, because a mistake reaches a patient.
Six questions. If you hesitated on more than one, the model was never your risk. Your blind spot was.
Why I am not writing this from the outside
I spent years building and scaling location and identity infrastructure inside hospital systems, first as a co-founder at Bluvision, later leading IoT and healthcare product strategy at HID Global following its acquisition of the company. Nearly every project that failed to scale past a pilot failed for the same reason CAG is describing now. The technology worked. The organization had not built the scaffolding required to trust it running unsupervised. Real time location data, device identity, and access control all hit the same wall agentic AI is approaching now, at greater speed and higher stakes.
The takeaway
The health systems that come out ahead over the next three years will not be the ones running the most pilots. They will be the ones willing to fail the six question test today, admit what it exposes, and only then let an agent operate on its own. CAG already proved this discipline works outside healthcare, at more than sixty percent rejection, in an industry with even less tolerance for error than yours. Healthcare does not need to reinvent that discipline. It needs to stop being surprised that it is required.