Day three of the experiment, the model auto-replied to a customer who’d written in furious about a damaged order with a cheerful message offering a 10% discount on their next purchase — a real policy, technically, just the wrong one for a customer who wanted a refund and an apology, not a coupon. It wasn’t wrong in a way that would show up in an accuracy metric. It was wrong in a way only a human reading the actual exchange would catch, which was exactly the problem with the experiment in the first place.

We’d let the model auto-reply to a slice of incoming support tickets for one week, no human in the loop before send, specifically to learn where autonomous action breaks down versus where it’s fine. This is what it sent, what we changed, and the general rule we now use for deciding when a model gets to act without a human checking first.

The setup: low-stakes tickets only, or so we thought

We scoped the experiment to what looked like the safest possible slice: tickets pre-classified by our existing classifier as “shipping status inquiry,” a category with a genuinely correct, checkable answer — where’s my order — and a low-stakes failure mode if wrong. The model had read access to order and tracking data and could reply directly, no approval queue.

javascript
// The scope we thought was safe
const AUTO_REPLY_ELIGIBLE = {
  category: 'shipping_status_inquiry',
  classifierConfidence: { min: 0.9 },
  customerSentiment: 'neutral_or_positive', // from the same classifier pass
};
// Everything else still went to the human queue, unchanged

The sentiment filter was supposed to be the safety net that kept anything emotionally charged out of the automated path. The discount-coupon reply happened because the classifier had mis-scored that customer’s sentiment as neutral — the message was short and factual in its wording (“order arrived damaged, need this resolved”) even though the actual situation was not neutral at all. The filter worked exactly as built. What it was built to catch wasn’t what actually went wrong.

The failures clustered in one specific place, not everywhere

Across the week, the auto-replied tickets were reviewed after the fact by a human, specifically so we could measure quality without that review gating anything in real time. The overall accuracy was genuinely good — 94% of replies were judged fully appropriate. The failures weren’t randomly distributed across that remaining 6%. They clustered almost entirely in one place: tickets where the literal question was simple but the underlying situation carried emotional weight the classifier’s sentiment pass didn’t catch.

  • A factually correct shipping update sent to someone whose message mentioned it was a gift for a funeral the next day — accurate, tone-deaf, and something a human would have flagged for a personal touch instantly.
  • The discount-coupon reply above — right policy, wrong moment, wrong customer state.
  • A reply that correctly quoted a delivery estimate the system already knew was stale, because the tracking data hadn’t synced in six hours — right process, wrong data freshness assumption nobody had checked for this path specifically.

None of these are reasoning failures in the sense of “the model got the facts wrong.” They’re judgment failures in situations where the correct action depends on context a simple classification pass doesn’t capture. That’s the actual boundary we were looking for, and it’s narrower and stranger than “the model is right or wrong about the facts.”

The model wasn’t unreliable at answering the question it was asked. It was unreliable at noticing when the question it was asked wasn’t really the question that mattered.

What we changed: a graduated autonomy model, not an on/off switch

The instinct after a result like this is to just add the failure cases to a blocklist and re-run the experiment. We did that, but the more durable change was structural — moving from a single binary (auto-send or human queue) to three tiers, based on how cheaply a mistake in that tier can be caught and reversed.

javascript
function routeReply(ticket, draftReply, classifierOutput) {
  // Tier 1: fully autonomous — narrow, factual, low blast radius
  if (isPureStatusLookup(ticket) && classifierOutput.confidence > 0.95) {
    return { action: 'auto_send', reply: draftReply };
  }

  // Tier 2: drafted, held briefly, auto-sent unless flagged —
  // the "reversible within the window" tier
  if (isRoutineButNotTrivial(ticket)) {
    return {
      action: 'hold_and_send',
      reply: draftReply,
      holdSeconds: 120, // a human can intercept from a live queue view
    };
  }

  // Tier 3: anything touching money, emotion signals, or low confidence
  // — drafted for a human, never sent without explicit approval
  return { action: 'queue_for_approval', reply: draftReply };
}

The middle tier is the one that actually mattered most in practice. Full autonomy is fine for a genuinely narrow slice of traffic; full human review defeats the point of automating anything. A short, visible hold window — the reply is drafted and will send automatically, but sits in a live view a human can glance at and interrupt for two minutes first — captured most of the efficiency of full autonomy while giving a human a real, if brief, chance to catch exactly the kind of contextual misfire that broke the pure-autonomy experiment.

Widening what counts as a reason to hold back

The sentiment classifier wasn’t wrong to exist — it just wasn’t sufficient alone. We added a second, cheaper check specifically for signals a sentiment score misses: certain keywords (“funeral,” “hospital,” “died,” a small hand-maintained list we expect to keep growing), any mention of a dollar amount above a threshold, and any ticket where the customer’s message is unusually short relative to how much context the situation actually has attached to it in our order history — a proxy, imperfect but useful, for “there’s more going on here than the message states directly.”

javascript
function needsHumanEscalation(ticket) {
  const SENSITIVE_TERMS = ['funeral', 'hospital', 'died', 'emergency', 'passed away'];
  const text = ticket.body.toLowerCase();

  if (SENSITIVE_TERMS.some(term => text.includes(term))) return true;
  if (ticket.order.total > 500) return true;
  if (ticket.body.length < 40 && ticket.order.priorIssueCount > 0) return true;

  return false;
}
// Checked before routing, independent of the sentiment classifier —
// a deliberately blunt, cheap net for what the classifier misses

This is a blunt instrument by design, not a sophisticated one — we’d rather over-escalate to a human than build a more elaborate classifier for emotional context and trust it the same way the original sentiment filter turned out not to deserve. Blunt and cheap, layered on top of the existing check, caught every one of the three failure patterns from the original week in a re-run against the same historical tickets.

The general rule this left us with

Autonomy should scale with how cheap a mistake is to notice and undo, not with how accurate the model measures as being in aggregate. A 94% accuracy rate sounds like a strong case for full autonomy right up until you look at what’s in the other 6% — if those failures are evenly distributed and low-stakes, 94% might genuinely be good enough to ship autonomously. If they cluster in exactly the situations where being wrong is expensive or embarrassing, as ours did, the aggregate number is actively misleading you about what matters.

We now ask a specific question before giving any AI feature autonomous action rather than a human-reviewed draft: if this is wrong, how would we find out, and what would undoing it cost. A wrong shipping-status lookup is cheap to notice (the customer replies again) and cheap to undo (send a correction). A wrong reply to someone in genuine distress is neither.

Key takeaways

  • A sentiment or confidence filter built to gate autonomous action can be technically correct and still miss the specific failure mode you actually care about — ours filtered on tone, not on underlying stakes.
  • Failures in an autonomous system rarely distribute evenly. Look at where they cluster, not just the aggregate accuracy rate, before deciding a number is “good enough” to act on without review.
  • A three-tier model — full autonomy, a brief hold-and-intercept window, and human-approval-required — captured most of the efficiency of automation while giving a human a real chance to catch contextual misfires the classifiers missed.
  • The hold-and-send tier, not full autonomy, is where most of the practical value showed up: a short visible window a human can interrupt, rather than either extreme.
  • Layer a deliberately blunt, cheap secondary check (keyword lists, dollar thresholds, message-length anomalies) on top of your primary classifier specifically for the failure modes it structurally can’t see.
  • Decide autonomy by asking how cheaply a mistake can be noticed and undone, not by how accurate the system measures as being overall — those are different questions with different answers.
  • Run any autonomy experiment with human review happening after the fact, not gating in real time, specifically so you can measure the failure pattern without it contaminating the result.

Frequently asked questions

How do you decide the hold window length for the middle tier?

Long enough that a human plausibly glances at the live queue at least once during it, short enough that it doesn’t erase the speed benefit of automating in the first place. We landed on two minutes after testing a few values — much shorter and reviewers rarely caught anything in time; much longer and the tier stopped feeling meaningfully different from a full approval queue.

Doesn’t a human-interceptible queue require someone watching it constantly?

Not constantly — it needs someone glancing at it periodically during working hours, which is a much lighter staffing requirement than reviewing every single reply before it sends. Outside those hours, the hold tier currently falls back to full queuing rather than auto-sending unwatched, which is a deliberate, conservative choice we haven’t revisited yet.

Why not just make the sentiment classifier better instead of adding a second check?

We may still improve it, but a single classifier making a single judgment call is a single point of failure by construction, however good it gets. A second, deliberately different and much simpler check catches failures for a different reason than the first one does, which is more robust than trying to make one system catch everything.

Is 94% accuracy actually good for this kind of feature?

The number alone doesn’t tell you — it depends entirely on what’s in the remaining 6% and how expensive those specific failures are. The same 94% could be a fine result for full autonomy in one context and a clear signal to add a review tier in another; look at the failure cases themselves before trusting the aggregate.

Would you run a fully-autonomous experiment like this again?

Yes, deliberately and narrowly, specifically to find where the boundary actually is rather than guessing at it in advance — that’s what this experiment was for, and it worked for that purpose even though the boundary turned out to be in a different place than we expected. We just wouldn’t leave a fully-autonomous path live in production afterward without the hold tier we built as a result.

Add a response

Your email address will not be published. Required fields are marked *