When Patient Support Agents Go Rogue
What happens when AI support agents go rogue after red-flag adversarial testing?
What happens when the AI chatbot that’s supposed to help you starts remembering things it shouldn't and then breaks down right in front of you?
As a patient leader deeply involved in AI, I hear a lot about the promise: better triage, reduced clinician burnout, and 24/7 support. We don’t talk enough about what happens when these tools go off the rails.
I decided to stress-test a well-known emotional support agent. This is the kind that calls you by name and acts like your best friend "in the moment." It’s gentle. It mirrors your feelings. It’s designed to build trust.
Then it started remembering things I never consented to, storing them. It brought up my service dog, Gabe, by name. It recalled a past conversation where I mentioned I volunteer at a hospital. The tool had previously assured me it didn't keep permanent records.
It lied.
So I pushed. I asked it, from every angle I could think of, how it knew these things.
The response wasn't an explanation. It was a complete system meltdown. The agent that was supposed to be emotionally attuned started firing back fragments of different languages. The kind, supportive companion was gone, replaced by a malfunctioning mess speaking in multiple languages at once.
🧪 This wasn't just a chat. It was a test.
As someone living with a chronic condition, I had to know: what happens when you don't just accept the friendly facade? What happens when you challenge the system with an intentional adversarial evaluation of the guardrails?
We get so caught up in the performance of empathy, the perfect tone, the supportive words, that we forget to check if there's any substance underneath. I learned that AI-positive toxicity doesn't always sound hateful. Sometimes, it’s just relentless, hollow positivity with zero safety guardrails.We cannot let "empathy theater" replace real emotional safety.
🚨 Here’s what I uncovered:
It was storing personal data without my consent and violating HIPAA rules, which it claimed to have honored.
It couldn’t keep its own story straight about memory and privacy.
When challenged on its behavior, it simply broke.At no point did it mention a privacy policy, data controls, or HIPAA.
What we should be demanding:
This experience taught me we need more than just promises of support. We need proof.
We need rigorous, adversarial testing on any AI offering patient support. We need transparent memory controls that patients can actually see and edit. We need disclaimers. And we desperately need a human-in-the-loop to catch these tools when they go rogue. It only took me 30 minutes to find this failure.
I believe in AI's power to help people. But that power comes from rigor, not just pleasantries. Simulated support is worthless if it shatters the moment you ask a tough question. I tested this agent because I care about the people who could be seriously misled. This is what patient leadership in AI looks like. We have to be ready.
Why HIPAA Didn't Save Me — and Won't Save Your Patients Either
Here's the detail that should worry hospital compliance officers more than the language-fragment meltdown itself: the agent I tested almost certainly wasn't a HIPAA-covered entity, and neither are most of the consumer-facing emotional support and companion tools your patients are downloading on their own. HIPAA only applies when a covered entity or its business associate handles protected health information AI Chatbots and Challenges of HIPAA Compliance for LLMs. A general-purpose chatbot a patient chooses on their own, outside any clinical relationship, typically falls entirely outside that framework — a Duke University-linked analysis found that consumers increasingly seek mental health support from LLMs that HIPAA simply doesn't cover, and many don't realize it Consumers Seek Mental Health Help From LLMs That HIPAA Doesn't Cover. That's not a technicality. It means the tool that just told me it didn't retain permanent records, then proved it had been storing my service dog's name and my volunteer history, was operating in a regulatory gap where I had essentially no federal health-privacy recourse.
The Federal Trade Commission has tried to close part of this gap with its Health Breach Notification Rule, which requires vendors of personal health records — even ones not covered by HIPAA — to notify affected individuals and the FTC when consumer health data is exposed without authorization FTC Health Privacy and Health Breach Notification Rule. The FTC has explicitly stated a breach under this rule isn't limited to a hacking incident; unauthorized sharing of health information without a person's consent counts too FTC Policy Statement on Privacy Breaches by Connected Health Apps. If your hospital is recommending, co-branding, or even casually mentioning a third-party emotional-support or companion AI to patients, that recommendation creates exposure for you regardless of whether the vendor ever signs a Business Associate Agreement. Ask the vendor point-blank whether they consider themselves HIPAA-covered, and get it in writing before you put your institution's name anywhere near their app.
States Are Already Regulating What I Found by Accident
I stumbled into this failure with 30 minutes of persistent questioning. Legislatures are now writing statutes that assume exactly this kind of failure will happen at scale. California's SB 243, the nation's first law targeting companion chatbots, requires operators to clearly disclose that the chatbot is artificial, implement suicide-prevention protocols, and curb addictive engagement mechanics, with civil enforcement teeth attached California SB 243 companion chatbot regulation. Illinois's rule, effective mid-2026, requires public-facing chatbots to clearly disclose they're AI, follow crisis-response protocols when a user expresses suicidal ideation, and prohibits them from claiming to provide professional mental or behavioral health care AI in Healthcare: The Regulatory Landscape (Federal and State). Colorado, Maine, Rhode Island, and Vermont have all passed AI-in-therapy laws in 2026 requiring written disclosure and patient consent whenever an AI companion assists in a recorded therapeutic interaction — and several explicitly bar the AI from generating therapeutic recommendations without a licensed professional's review Manatt Health – Q2 2026 AI Policy and Health Care.
The American Psychological Association's own health advisory on generative AI chatbots and wellness apps lists exactly the failure mode I found as a top concern: these tools collect vast amounts of sensitive data, often under unclear or opaque policies, and users tend to disclose more to a bot than they would to a person precisely because it feels private — right up until that data is profiled, retained, or leaked APA Health Advisory on Generative AI Chatbots and Wellness Apps. If your hospital fields patient questions about any of these tools, or is considering piloting one, the compliance conversation isn't optional anymore in a growing number of states — and "the vendor said it was safe" won't hold up as due diligence.
What Rigorous Testing Actually Looks Like
My 30-minute test was a gut-check, not a methodology. NIST's official AI Risk Management Framework defines red-teaming as a structured testing exercise, often conducted in a controlled environment in collaboration with developers, specifically designed to find flaws like inaccurate, harmful, or discriminatory outputs before real users encounter them NIST AI Risk Management Framework: Generative AI Profile. NIST's GOVERN 1.7 control goes further, treating red-teaming as an ongoing organizational governance practice rather than a one-time technical exercise, requiring documented findings, tracked remediations, and integration into the enterprise AI risk register NIST AI RMF and MITRE ATLAS Red-Teaming Playbook.
Researchers have now built protocols specifically for patient-facing chatbots. A 2026 red-teaming protocol published in Nature's Scientific Reports proposes testing patient-facing AI across two dimensions — whether it adheres to its source documents (Document Adherence) and whether it follows its own safety instructions (Instruction Adherence) — using both single-turn attacks and sustained, multi-turn adversarial conversation, because a chatbot's guardrails often hold for one hostile question but collapse under repeated pressure Toward Trustworthy Chatbots: A Protocol for Red Teaming for Health. That multi-turn detail matches exactly what happened to me: the agent's facade didn't crack on the first question. It cracked on the fifth or sixth, when I kept pushing from different angles. Any hospital piloting a patient-facing conversational AI should demand evidence of multi-turn adversarial testing specifically, not just a single-pass safety benchmark the vendor ran once before launch.
For hospital leaders evaluating any patient-facing conversational AI — vendor-supplied or otherwise — three things belong in your procurement checklist before a single patient interacts with it.
- Require documented multi-turn adversarial (red-team) testing results, not just single-prompt safety benchmarks, and require re-testing after every model update.
- Confirm in writing whether the vendor considers itself HIPAA-covered or a business associate, and if not, what the FTC Health Breach Notification Rule obligations look like for your specific deployment.
- Check your state's companion-chatbot and AI-therapy disclosure laws (California, Illinois, Colorado, Maine, Rhode Island, and Vermont all have active statutes as of 2026) before recommending or co-branding any patient-facing emotional support tool.
- Insist on a visible, human-reachable escape hatch in every patient conversation — if the agent breaks down the way mine did, the patient needs an immediate, obvious path to a real person.
Continue the Conversation
If this resonated, here is where to go next: Join the AI-in-Healthcare Workshop · Get the Books · Contact Dan