Putting AI to Work: What Patient Safety Leaders Need to Know About AI Agents


Some of the most important work in quality and safety is also among the hardest to streamline. Incident narratives, patient complaints, recalls, and other safety signals require teams to interpret context, recognize patterns, and decide when information warrants action. Yet much of the information needed to make those decisions is unstructured, leaving a relatively small group of experienced quality, safety, risk, and compliance professionals to sift through it, connect the dots, and determine what needs attention.
A new generation of artificial intelligence (AI) can finally reach that work. But before a safety leader hands any of it to a machine, one distinction matters more than any other: the difference between an AI agent and an agentic workflow. That single line determines how much risk an organization is taking on and how much governance it owes in return.
Why this conversation is happening now
It helps to know how we got here. In a widely cited paper, physicians from Google described three epochs of AI in health care. The first was rules-based: if a lab value crossed a threshold, an alert fired. These systems still run most hospitals, but they act only on structured data someone has already coded. The second epoch, machine learning, tapped unstructured data at scale and produced tools like sepsis prediction models, but often as a “black box” that could not explain its reasoning.
The third epoch is generative AI, and it changes the equation. For the first time, software can read an incident report the way a safety event reviewer would, reason across structured and unstructured data together, plan a sequence of steps, and, critically, show its work for human review. The question is no longer whether the technology can do useful work, but rather how much autonomy to grant it, where to deploy it first, and how to govern it.
The distinction that determines your risk
An AI agent is a tool a human built, scoped, and placed guardrails around. It acts on your behalf within limits you defined, and you can see what it did. An agent that drafts an acknowledgment letter for a patient grievance and routes it to the patient advocate for approval is working this way. It is autonomous in execution but not in judgment, because a human set the boundaries in advance.
An agentic workflow is different. The system reasons about its own authority in real time. It decides what matters, what warrants escalation, and what it can handle alone, interpreting situations and choosing actions without waiting for a person. This is where the significant risk lives, and it is also where the significant value lives. The two arrive together and cannot be separated.
Here is the honest tension worth naming. If you constrain an agentic system so tightly that it never makes a judgment call, you have not built an agentic system. You have built a rule engine with a language model attached, and you are back at the ceiling that limited the first epoch. Genuine autonomy introduces new kinds of risk. It is also what frees skilled professionals to return to the judgment work only they can do.
One boundary should never be crossed: no system should set its own limits. Those limits belong to the humans who built and govern it.
Where agents earn their keep — and where humans stay in control
The uses with the greatest potential do not sit in the corners of clinical practice. They sit in the daily work of quality, safety, risk, and compliance, where volume is high, deadlines are real, and most of the information is unstructured. Consider four familiar workflows:
Incident reporting. A nurse files a medication-error narrative at 2 a.m. An agent can read it immediately, classify severity, check the drug against the high-alert list, and recognize that this is the third event involving the same infusion pump model, flagging the pattern with supporting evidence. The safety event reviewer arrives to a triaged list instead of a stack of raw reports.
Recall and safety alert management. Notices from the U.S. Food and Drug Administration (FDA), manufacturers, and group purchasing organizations arrive as unstructured text. An agent can extract affected products and lot numbers, match them against inventory and pharmacy stock, and help answer the urgent question — was any patient exposed? — in hours rather than weeks.
Patient complaint and grievance management. An agent can read each complaint as it arrives, distinguish a complaint from a grievance, start the regulatory response clock, route the case, and draft an acknowledgment, while surfacing links to safety events on the same unit or provider.
Quality measure abstraction. An agent can extract from the documentation itself, flag a measure failure in near real time, and route it to the responsible service line with the specific gap identified — the difference between correction and month-end explanation.
Notice the common thread. In each example, the agent handles routing, drafting, monitoring, and pattern detection, where errors can be caught and undone. A human stays in control wherever the action crosses into patient care, external reporting, or anything irreversible. This line is the heart of safe adoption, and it is not drawn by any vendor. Your organization draws it, documents it before deployment, develops an implementation plan, and changes it only through a formal change process.
Trust is earned, not granted
Leaders will ask the obvious question: how do we know the system is making the right call? The honest answer is that trust is earned through evidence, much the way a team learns to trust a new hire. With agents, AI can automate a task and remove human involvement together. A sound practice is to run the system in recommend-only mode and measure how often its proposed action matches what your people actually do, then expand its scope only as that agreement rate earns confidence. A system that cannot show its work does not belong in a healthcare workflow.
The new duty to monitor
This is a total-systems problem, not a software purchase. Before AI, organizations governed vendor software as a one-time evaluation: the tool did the same thing on day 500 that it did on day one. That world is gone. An AI system changes as models drift, data shifts, and workflows evolve around it. Governing AI is not a project with an end date; it is an ongoing operational duty that lasts as long as the system runs.
Effective governance rests on a few durable elements: risk-based classification of every deployment, explicit written boundaries, audit trails and plain-language explainability, continuous monitoring for drift, and a deliberate effort to preserve human capability so staff skills do not atrophy. One difference from managing people is worth underscoring: when an employee makes a mistake, it is usually contained; when an agent makes one, it can repeat across every case until someone catches it. That is why performance review here must be continuous, not annual.
The regulatory landscape is converging on the same expectations. The NIST AI Risk Management Framework, the Office of the National Coordinator’s HTI-1 transparency requirements, and the Joint Commission and Coalition for Health AI guidance on responsible AI use all point toward defined accountability, monitored performance, and human oversight proportionate to risk.
Key takeaways
Know which one you are deploying. An agent acts within human-set guardrails; an agentic workflow reasons about its own authority. The distinction sets your risk and your governance burden.
Draw the autonomy line yourself. Let agents handle routing, drafting, monitoring, and pattern detection; keep humans in control of patient care, external reporting, and anything irreversible.
Earn trust with evidence. Run in recommend-only mode, measure agreement with human judgment, and expand autonomy only as confidence grows.
Treat monitoring as permanent. Model drift makes AI governance an ongoing duty, not a one-time evaluation — the real limit on safe adoption is governance maturity, not the technology.
The automation ceiling is real, and this new generation of AI is how organizations move past it. The ones that get there first will not be the ones that moved fastest, but the ones that moved with discipline: drawing a clear line between assisted and autonomous action and governing these systems with the same care they would give a new member of the team.
References
Howell MD, Corrado GS, DeSalvo KB. Three epochs of artificial intelligence in health care. JAMA. 2024;331(3):242-244. doi:10.1001/jama.2023.25057
National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. January 2023.
Office of the National Coordinator for Health Information Technology. Health Data, Technology, and Interoperability: Certification Program Updates, Algorithm Transparency, and Information Sharing (HTI-1) Final Rule. 89 FR 1192. January 9, 2024.
The Joint Commission and Coalition for Health AI. Guidance on Responsible Use of AI in Healthcare. September 17, 2025.
AI disclosure: An AI tool was used to assist with outlining and drafting this article from an Ideagen white paper; all content was reviewed and edited by the contributing team.
As a Founding Strategic Partner, Ideagen is part of MPSC’s Council of Industry Partners—a community of organizations committed to advancing patient safety through collaboration. Learn more about becoming a Council partner.




