An AI assistant in a customer's inbox is judged on its worst answer, not its average one. Every setting worth understanding is a way of making the worst answer a handoff instead.
The assistant has two sources: a persona, which is how it speaks, and a knowledge base, which is what it knows. Without a knowledge base it answers from the persona alone — fine for tone, and guesswork for facts. It will sound like you while being wrong about your opening hours.
A knowledge base is your own content: an FAQ, your policies, the pages of your site that answer questions. The assistant draws on it and only on it for facts, which is what makes an answer checkable — every reply records which source it used, so a wrong one can be traced to the page that caused it and fixed there rather than in the assistant.
Every answer comes with how sure the assistant was, and the floor is the level below which it does not answer at all — it hands the conversation to you instead. At zero, the knowledge base's own threshold is in charge. Raise it when a wrong answer costs more than a slow one: a refund policy, a medical question, a price.
A high floor produces more handoffs and fewer mistakes; a low one the reverse. Neither is right in general. The activity log shows the confidence on every reply it gave, which is the evidence for where to set it — look at the answers near the floor and ask whether you would have wanted them sent.
The assistant stops and a person takes over in four situations: it was not sure enough, the customer is upset, the customer looks like a hot lead, or a keyword you chose appeared. In every one the conversation arrives in your inbox already in full, with a holding message sent to the customer so they are not left waiting in silence.
Two channel rules shape this. On WhatsApp the assistant cannot reply outside the 24-hour window any more than you can — a handoff there waits for the customer to write again, or becomes an approved template. On Messenger the human agent tag, which extends the window to seven days, applies to a person's reply and never to the assistant's; the product does not let the assistant use it, and that is deliberate.
A conversation that goes in circles is worse than a handoff, so the turn limit caps how many times the assistant may reply in one conversation before handing over regardless of confidence. Zero means no cap. A small number — five, eight — is usually right: a question that has not been resolved in eight exchanges is not going to be resolved in twelve.
Everything above can be tested before the assistant is switched on. Send it the questions you actually get, read what it says and how sure it was, and mark the bad answers — each one feeds back into the knowledge base rather than into a setting. Ten minutes of this finds the gaps that would otherwise be found by a customer.
Each reply costs credits, and bigger models cost more per reply, so a smaller model goes further on the same allowance. The AI page shows what has been used; the trade-off is the same one as the confidence floor, and the activity log is again where the evidence lives.
It is the level of certainty below which the assistant does not answer and hands the conversation to you instead. Zero leaves the knowledge base's own threshold in charge. Raise it when a wrong answer costs more than a slow one.
The conversation arrives in your inbox in full, a holding message goes to the customer so they are not waiting in silence, and the assistant stops replying in that thread. Handoff happens on low confidence, an upset customer, a hot lead, or a keyword you chose.
No. The window applies to every message from your number, automated or not. Outside it only an approved template will send, so a handoff there waits for the customer to write again or is answered with a template.
Not to run, but without one it answers from its persona alone — good for tone and guesswork for facts. A knowledge base built from your own FAQ, policies and pages is what makes its answers checkable and its mistakes fixable at the source.