How Does Human-in-the-Loop AI Work in Surgical Patient Care?
Human-in-the-loop is not a single approval button. It is an operating model that defines what the AI may do, what it must not do and where people remain accountable.
Human oversight begins before the first patient message
A responsible AI workflow is designed by people before it is used by patients. The surgical program decides which use cases are appropriate, which content is approved, what information may be collected, which escalation criteria apply and who is responsible for review.
NIST’s AI Risk Management Framework describes AI risk management as an ongoing organizational process rather than a one-time technical checklist. The framework emphasizes governance, context, measurement and risk management across the AI lifecycle. [5]
For a patient-facing surgical companion, that means human oversight must exist in the knowledge sources, workflow design, access controls, review process and ongoing monitoring.
Layer 1: Program-approved knowledge boundaries
The customer uploads its own surgical protocols, care guide and internal templates. These materials define the information Waivs may use for routine patient communication. An approved knowledge boundary is more useful than a broad promise that the AI “knows surgery.”
Clinical leaders should review the source materials before implementation and approve changes before updated content becomes active. Outdated, contradictory or incomplete materials should be resolved rather than passed directly into an automated workflow.
When a patient asks a question outside the approved content, the system should not improvise a clinical answer. It should identify the boundary and move the interaction to human review.
Layer 2: Context without autonomous judgment
A surgical workflow can use context such as procedure type, pre-op or post-op stage, days since surgery and previous responses. These variables help select the relevant approved questions, education and escalation criteria.
Context is not the same as diagnosis. The AI may determine which configured workflow applies, but a clinician interprets the patient’s condition and makes treatment decisions. CPSO guidance similarly states that AI is intended to complement clinical care and that physicians remain accountable for its use. [4]
This distinction should be visible in patient and staff materials: the system identifies matches and organizes information; the care team makes clinical judgments.
Layer 3: Predefined escalation criteria
The program defines the responses, patterns or conditions that should prompt review. Different criteria may apply by procedure, stage or recovery timing. The system identifies when a patient response matches those predefined criteria and surfaces the interaction to the appropriate team member.
An escalation is not a diagnosis and should not be described as one. It is a workflow signal that tells the program a human should review the information.
Teams should test the criteria against routine, ambiguous, incomplete and concerning scenarios. They should also review false positives, missed matches and changes in clinical protocols over time.
Layer 4: Role-based permissions and accountable review
Administrators determine who can see which information on the administrative side. Surgeons, nurses, coordinators, dietitians and other team members may receive different access based on their responsibilities.
Permission design should follow the minimum necessary principle: people should have the access required to perform their role without exposing information unnecessarily. Privacy and security risk analysis should consider how data are created, accessed, transmitted, storedand reviewed. HHS describes risk analysis as foundational to selecting safeguards for electronic protected health information. [6]
The implementation should make ownership explicit. Every escalated workflow needs a responsible role, an expected review process and a way to audit what happened.
What happens when the AI is uncertain?
Uncertainty is not a failure if the workflow is designed to handle it. The safer response is to remain within approved content, avoid an unsupported answer and route the interaction for human review.
Programs should test how the system responds to vague wording, multiple concerns in one message, contradictory answers, missing information and requests that fall outside the companion’s purpose.
A trustworthy implementation does not claim that hallucinations or errors are impossible. It creates controls that reduce risk, expose uncertainty and keep people responsible for the final decision.
Questions leaders should ask before launch
Which protocols and templates are approved for patient use? Who can change them? What context determines the workflow? Which responses trigger review? What happens when the answer is uncertain? Who can see the interaction? Who owns the escalation queue? How are changes tested? How will the program review performance and incidents?
These questions turn “human-in-the-loop” from a marketing phrase into an operating model. If the answers are unclear, the program is not ready to scale the workflow.
The core principle
AI can extend the reach of approved communication and make patient-reported information easier to organize. It should not transfer clinical accountability away from the surgeon or care team. The strongest implementation is one in which the boundaries, escalation logic and human responsibilities are understandable to everyone using it.
Sources and further reading
- College of Physicians and Surgeons of Ontario. Using Artificial Intelligence in Clinical Practice. Updated August 2025. https://www.cpso.on.ca/Physicians/Policies-Guidance/Advice-to-the-Profession/Using-Artificial-Intelligence-in-Clinical-Practice
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. 2023. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10
- U.S. Department of Health and Human Services, Office for Civil Rights. Guidance on Risk Analysis under the HIPAA Security Rule. https://www.hhs.gov/hipaa/for-professionals/security/guidance/guidance-risk-analysis/index.html
Related Articles
24/7 Patient Messaging Is Not the Same as 24/7 Clinical Coverage
Patients can need guidance outside office hours. Giving them a channel that is always available is valuable, but what does availability mean in the context of AI? Patient Messaging Availability vs. Clinical Coverage 24/7 patient messaging means a patient can send a message or respond to an automated check-in at any time. The system can receive the interaction, provide […]
How AI Fits into a Surgical Workflow – Save Coordinators Time for Patients Who Need It
The value of AI in a surgical workflow is not that it replaces judgment. It is that it separates repeatable communication from work that requires a clinician or coordinator. Patient Competition for Attention in your Inbox – High and Low Priority A patient asking which diet stage comes next and a patient describing a possible warning sign may reach […]
Three Places Bariatric Patients Disengage Before and After Surgery
Patient disengagement is common across surgical practices. It usually develops through small missed steps, unanswered questions and fading contact across a long bariatric care journey that can feel overwhelming to patients. Patient leakage is usually a pathway problem, not a single missed call Bariatric programs often use the term patient leakage to describe patients who stop progressing, disengage from follow-up or leave the care […]