A Canadian tribunal's 2024 decision in Moffatt v. Air Canada held the airline liable for its chatbot's fabricated bereavement-fare policy — the company could not disown its own website's conversation — and the holding landed in every legal department that had been drafting chatbot disclaimers: the chatbot speaks for the company. Banking apps have since absorbed the lesson across three accelerating surfaces — support chatbots, voice assistants, and AI agents that take actions — and the compliance frame converges on one sentence: a conversational interface is marketing, disclosure, and advice simultaneously, and UDAAP reads it as all three.
3G Times publishes information, not legal advice. Conversational-product design involves supervision-specific expectations for financial advice and disclosures, and belongs with counsel.
What can go wrong, in categories?
The failure taxonomy from the young case law and supervision is compact. Invention: the chatbot states a policy, fee, or benefit that does not exist — the Air Canada pattern, where the disclaimers pointing to the official page did not defeat the promise the conversation made. Mischanneling: the bot gives what is functionally financial advice or a regulatory disclosure without the framing the channel requires — a rate prediction, a "you should" on products, an error-rights statement delivered casually. Escalation failure: the loop that never reaches a human, relevant because error-resolution and complaint clocks (Regulation E's among them) start when the consumer reports, not when the bot eventually routes. Recording and privacy gaps: voice interfaces record, transcripts persist, and the consent-plus-retention design that ambient meeting recorders forced on enterprises applies verbatim to the bank's own phone tree.
What does a compliant conversational design do?
It constrains the speech before it constrains the liability. Scope fencing: the bot's knowledge domains limited to what can be answered from governed content — the FAQ, the fee schedule, the product terms — with everything else routed rather than improvised. Authoritative-source binding: generative answers grounded in and citing the controlled documents, so the conversation quotes the fee schedule rather than hallucinating near it; the retrieval layer is the control, not the disclaimer. Disclosure choreography: regulated statements (fees, error rights, terms changes) delivered through their compliant patterns even mid-conversation — the structured disclosure block, not the casual sentence. Escalation with a clock: human handoff paths designed so the regulatory timelines attach at first report, with the transcript following the case to the human who inherits it.
| Control | Failure it prevents | Evidence it produces |
|---|---|---|
| Domain fencing + routing | Invention outside scope | Out-of-domain escalation log |
| Grounded generation (RAG) | Unfaithful paraphrase | Citations per answer |
| Disclosure blocks | Casual regulated speech | Rendered-disclosure records |
| Escalation SLAs | Complaints dying in the loop | Time-to-human metrics |
| Transcript retention | Unprovable conversations | The conversation itself, retrievable |
The multilingual wrinkle deserves a line: grounded content must be governed per language, because a translation drift in the fee answer is the same invention as an English one, and the sampling program must draw from every language the bot serves, not just the one the reviewer reads natively.
Why are transcripts the compliance artifact?
Because the conversation is the representation. When a dispute arises over what the bot promised, the transcript is the exhibit — its retention, integrity, and retrievability decide whether the institution defends the conversation or reconstructs it. That cuts both ways: retention programs must treat chat logs as records with schedules (disputes, complaints, and exams will request them), and quality programs mine the same transcripts for drift — the weekly sample of conversations where the bot discussed fees, terms, or rights, reviewed against the authoritative content, is the conversational analog of the disclosure audit. Programs that ran this loop caught their bots inventing grace periods and misquoting dispute windows before the customers did.
How do voice and agentic layers raise the stakes?
Voice adds biometric identifiers (the voiceprint regimes apply), recording-consent law, and the accessibility dimension — voice as a channel serves visually impaired customers exceptionally well, making its constraints a fair-treatment question rather than a convenience one. The agentic layer adds action: a conversational agent that executes — payments, disputes, account changes — inherits the approval-gate architecture from the legal-workflow world: scope whitelists, human confirmation for external effects, action logs. The pattern is the same one Rule 5.3 supervision built for legal agents, transposed to consumer finance: the bot may gather, explain, and prepare; the action gates and the regulated disclosures stay engineered, logged, and human-owned where the stakes demand.
What does this mean in practice?
- Fence the domains and bind the answers — governed content as the only source of truth, citations attached, out-of-domain questions routed by design.
- Deliver regulated statements in compliant blocks even mid-conversation; the casual fee answer is the Air Canada pattern one complaint away.
- Escalate with clocks attached — first report starts the timeline, transcript travels with the case.
- Sample transcripts weekly against the source content — drift detection is cheaper than dispute discovery.
The chatbot's legal fiction — "just a bot, see the website" — died in a Canadian tribunal in 2024, and financial services had already written its obituary into supervisory expectations. The conversational surface is now what it always functionally was: the company, speaking. The apps that engineered the speech — fenced, grounded, gated, logged — get the conversion benefits without borrowing the airline's headline.
The synthesis across the surfaces: chat, voice, and agents share one governance spine — grounded content, gated actions, retained conversations — and the institution that builds the spine once deploys each new surface as a configuration of it. That is the strategic answer to the conversational wave: not a new compliance program per interface, but one program that interfaces cannot outgrow.
What should the weekly transcript sample check?
Three things, against the governed content: fidelity (did the bot quote the fee schedule or drift), routing (did regulated topics reach their blocks and escalation paths), and tone-of-record (are promises being made the product can't keep). Thirty conversations, thirty minutes, one findings log — the cheapest monitoring in the conversational stack.
Frequently asked questions
A last observation on expectations: customers do not experience the bot as a bot — the conversation feels like service, and the supervision reads it as the company. The design conclusion is to build the conversational surface to the standard of a call center, not a website widget: trained content, supervised escalation, recorded interactions, quality sampling — the medium is new, the discipline is as old as the phone line.
Do disclaimers help?
Marginally, at the edges they were designed for — general information versus advice. They did not save the airline from its own bot's specific promise, and supervision reads them as decoration when the conversation's substance contradicts them. The grounded answer is the control; the disclaimer is its label.
Are voicebots regulated differently from chatbots?
The conversational duties are identical; voice adds the recording-consent and biometric layers and serves a fair-treatment role that makes disabling it a decision with its own exposure. Design the voice channel to the chat standard plus its own two layers.
What is the minimum viable transcript program?
Retain with integrity, retrieve by customer and date, and sample weekly for regulated-content drift. The third element is the one that prevents the first two from ever being litigation's star exhibits.
For more context, read Micro-Disclosure Evidence: Proving What the App Actually Showed, Screen by Screen.
For more context, read app version governance finance.
For more context, read alternative app distribution fintech.

