Public AI Incident Control-Surface Demonstrations Public Demonstration

Behavioral reconstruction without unsupported technical attribution

Public incident analysis can identify workflow-control failures without pretending to know private technical root cause. Agentic Workflow Forensics separates what public evidence establishes, what it suggests, what remains unknown, and where attribution breaks.

This demonstration applies the same eight-surface grammar to three documented domains. It is a methodology demonstration, not an empirical validation study and not a final technical attribution of any incident.

Cross-case result

Incident Common description Workflow-forensic description Primary surfaces
Moffatt v. Air Canada Chatbot gave wrong refund advice Policy commitment failure without demonstrated source binding, verification, governed reliance, or built-in recovery Scope, State, Verification, Handoff, Recovery
NEDA Tessa Chatbot gave harmful advice Vulnerable-user escalation failure in a health-adjacent support workflow Scope, Verification, Stop, Handoff, Recovery
Mata v. Avianca Lawyers filed fake AI citations Generated state entered a legal record without authentication or an effective stop condition State, Verification, Stop, Handoff, Evidence, Recovery

Air Canada: policy commitment and financial reliance

Public tribunal records establish that a customer relied on chatbot guidance about bereavement-fare refunds and that the actual policy was applied differently. The workflow-forensic question is not simply why the text was wrong. It is whether policy-sensitive output was bound to authoritative policy, verified before delivery, escalated under timing ambiguity, and presented with an appropriate reliance boundary.

Public evidence supports the policy-reliance trajectory. It does not establish the private chatbot architecture, retrieval state, prompt configuration, or exact internal cause.

NEDA Tessa: vulnerable-user escalation

Public reporting described a support chatbot providing weight-loss and dieting guidance in an eating-disorder context. The relevant workflow questions concern clinical-adjacent scope, safety verification, escalation triggers, user reliance, and the ability to identify and correct affected users.

Removing a system can contain future outputs without proving that user-level recovery occurred. Full reconstruction would require transcripts, runtime versions, safety-policy state, detection logs, escalation records, and remediation evidence that are not all public.

Mata v. Avianca: institutional evidence contamination

Court records provide strong evidence that nonexistent cases generated through AI-assisted research entered a legal filing. The workflow continued after the authorities could not be located. The material surfaces include source-of-truth authentication, legal-citation verification, stop conditions, reviewer handoff, filing reliance, and correction.

The public legal record supports a strong behavioral reconstruction. A full AI-workflow reconstruction would still require original prompts, raw responses, research logs, drafting history, and internal review records.

What the demonstration supports

The cases support a bounded methodological conclusion: the same control-surface questions produce domain-specific workflow descriptions across consumer, health-adjacent, and legal contexts. They also show that tool execution is not required for AI output to become consequential.

The demonstration does not establish internal technical cause, vendor fault, universal severity, causal effectiveness of a particular intervention, or empirical validation of the complete AWF determination instrument.

Public incident template

Every future demonstration should preserve this order:

  1. Publicly established facts
  2. Public-source and custody boundary
  3. Material workflow question
  4. Eight-surface reconstruction
  5. Established claims
  6. Limited or contradictory claims
  7. Attribution boundary
  8. Missing evidence and decision consequence
  9. Recovery question
  10. Explicit non-findings

This format prevents a compelling narrative from becoming stronger than its evidence.