Every support tool now says "AI-powered." Almost none of them say what that means in practice for a two-person store: What does the AI actually do? When does it act alone? What does an AI-handled ticket really cost? And what happens when it's wrong, in your name, to your customer?

This is the practical guide we wish existed when we started — vendor-neutral where it can be, and openly ours where it can't. (We build TaskBees; we'll flag every place that bias matters.)

What AI support genuinely does well in 2026

The honest capability list, from running it on our own store daily:

  • Drafting replies is solved. Given your policies and the customer's message, current models write replies you'd happily send — matching tone, answering multi-part questions, staying courteous at 2am.
  • Grounded answers are solved if the grounding exists. An AI connected to your real order data answers "where is my order?" with the actual carrier and tracking number. An AI without that connection writes confident filler. The difference is plumbing, not intelligence — see how we handle WISMO.
  • Triage is solved. Classifying urgency, spotting the angry message, flagging the legal-ish one — reliable and immediately useful.
  • Judgment is not solved. Refund exceptions, threats, edge cases your policies never contemplated, a customer whose third email contradicts their first — these still want a human. The right architecture assumes it.

The autonomy ladder — the decision that matters most

Every AI support product sits somewhere on this ladder. Where yours sits determines both your risk and your bill:

  1. Draft-assist. AI writes, human sends. Zero autonomous risk; you're the bottleneck, but a fast one — approving takes seconds, writing took minutes.
  2. Approval-first with earned autonomy. AI drafts everything; specific types of tickets unlock auto-send only after the AI proves itself on your own approval history — and you can revoke it. (This is TaskBees's model, so weigh our enthusiasm accordingly: we think it's the right trade for small stores, and we built a company on that opinion.)
  3. Autonomous with guardrails. AI resolves what it's confident about, escalates the rest. This is where most helpdesk AI add-ons live — and, notably, where per-resolution billing usually attaches: the vendor charges for each conversation the AI closes.
  4. Fully autonomous. No human in the loop. For a small brand whose reputation is personal, we'd argue nobody should start here, including with us.

The ladder maps to a simple rule: climb it with evidence, not hope. Start where mistakes are impossible, promote the AI per ticket-type as its record earns it.

The pricing-model traps

Three structures dominate, and the sticker price hides the difference:

  • Per-agent: classic helpdesk pricing. Cheap solo, scales with headcount — fine for small teams, but you're paying for coordination software to get the AI.
  • Per-resolution metering: the AI's work is a usage bill — commonly around a dollar per resolved conversation on top of your plan, per current market reporting. The perverse property: the better it works, the more you pay, and busy months bill highest. We've worked one real example in our Gorgias comparison.
  • Flat with AI included: one price, AI on every ticket. (Ours: $39/mo for 150 tickets.) The trade-off to check with any flat vendor — including us — is the ticket ceiling and what overflow costs.

Whichever you evaluate, ask the one question that cuts through all three: "If AI touches every one of my tickets next month, what is my total bill?"

The eight questions to ask any vendor (including us)

  1. What grounds the drafts? Your policies and live order data, or a generic model? Ask to see a draft's sources.
  2. Who approves sends, and can that change per ticket type?
  3. What does an AI-handled conversation cost, all-in? Plan fee + AI fee + overage, in writing.
  4. Is my data used to train shared models? (Ours isn't; make everyone answer it.)
  5. What's the migration cost — and the exit cost? Platform tools require both; inbox-native tools require neither.
  6. What happens on ambiguity? The right answer involves the words "asks the customer" or "escalates to you," never "picks the most likely order."
  7. How do I see what the AI got wrong? Edit-tracking and confidence reporting, or vibes?
  8. What channels are actually covered today? "Coming soon" is a roadmap, not a feature — judge on today's list. (Held to our own rule: our Shopify integration is coming soon; today we ground drafts in ShipStation and Shippo.)

A realistic first week

  • Day 1: connect the inbox (OAuth beats forwarding rules), upload policies/FAQ — the grounding matters more than the model.
  • Days 2–3: review every draft against what you'd have written. Edit freely; edits are training signal for tone.
  • Days 4–7: measure two numbers — percent of drafts sent as-is or lightly edited (quality), minutes per ticket before vs after (the point). A good setup sends most routine drafts nearly untouched within the week.
  • Then: consider autonomy only for the ticket types with a proven record, and only if your tool bounds it.

The honest limits

AI support won't save a store from bad policies (it will restate them faster), won't absorb a genuine crisis (a recall, a viral complaint — humans, immediately), and won't make judgment calls you haven't encoded. What it removes is the 80% of support that was always mechanical: the same five answers, personalized, at all hours. For a founder, that's not a productivity tweak — that's evenings.

Common questions

Is AI customer service safe for a small brand?

At autonomy levels 1–2, yes — the failure mode is a draft you didn't like and didn't send. The risk arrives with unsupervised sending, which is why we'd tell you to demand approval-first defaults from any vendor, us included.

Will customers know or care that AI helped?

Replies grounded in your data, sent in your voice, from your address, after your approval are your replies — the AI is how they got written fast. What customers notice is the reply time. Disclosure norms are yours to set; speed is what moves reviews.

Do I need this if I only get 20 emails a day?

Twenty emails is roughly an hour of context-switching daily — the question is what your hour is worth against a $39–99 tool. Below ~10 a day, honestly, maybe you don't need anything but templates and discipline.


The technology stopped being the hard part. The decisions that matter now are architectural — what grounds the AI, who approves it, and how it's billed. Choose those three deliberately and the vendor question mostly answers itself.