Small teams treat customer service metrics in one of two ways. Track nothing and run on vibes, or track everything the dashboard offers and stare at fourteen numbers that never change a decision. Both are ways of not deciding. I've watched hundreds of support conversations and nearly as many dashboards, and the teams that improve are not the ones with the most charts. They are the ones that picked a few numbers, understood exactly how each one lies, and met weekly to act on them. This post is that shortlist: the core four, the popular liars, and the new columns that appear when an AI agent answers first.
The core four
First response time
What it tells you: whether anyone is home. Visitors decide how reachable you are long before the answer arrives; the wait is the first answer they get.
How it lies: through averages. Fast daytime replies bury the overnight gap, and a decent median hides the person who waited two days and quietly left. It also invites tricks, like an instant auto-reply that technically "responds."
A sane way to collect it: measure against your stated promise, split into business hours and after-hours. If you promise a reply within a day, count the misses, not the average. Misses are what customers experience. After-hours deserves its own bucket: if nobody replies overnight, say so in the widget, and let that bucket measure whether the promised morning reply actually happened.
Resolution rate
What it tells you: how many conversations end with the question answered.
How it lies: "ended" is not "answered." In most tools silence counts as success, so a visitor who gave up looks identical to one who got what they needed.
A sane way to collect it: define resolution by a signal (an explicit "that helped," a thumbs up, no reopen within a few days) and keep the definition stable. Every definition has holes. A stable imperfect definition still shows the trend, and the trend is the part you can act on. Context matters when it moves, too: a drop the same week you launched a new product line usually means new questions arrived before the docs did, not that support got worse.
Contact rate
What it tells you: whether growth is creating support work. Count conversations per 100 orders for a store, per 100 active users for software. It is the only volume-independent number on the board, which makes it the one worth a spot on the wall.
How it lies: barely, which is why I like it. The main trap is changing the denominator mid-year and calling the movement a trend.
A sane way to collect it: one line in a spreadsheet, monthly. The math is small. Say last month was 2,000 orders and 80 conversations, a contact rate of 4. This month is 3,000 orders and 120 conversations. The inbox feels half again heavier and nothing is wrong; the rate held at 4 while the business grew. When the rate itself climbs, something upstream broke: confusing checkout copy, a missing shipping page, a settings screen nobody can find. Fixing those sources is the honest way to reduce support ticket volume, and a rising contact rate tells you where to look first.
CSAT, with caveats
What it tells you: how the people who bothered to answer felt. That is a real signal, just a narrow one.
How it lies: response bias. The furious reply. The delighted reply. The satisfied middle closes the tab. At low volume the score also swings on a handful of responses, so a bad week can be one person and a mispriced shipping label.
A sane way to collect it: one tap after resolution, nothing longer. At small scale, treat the comments as the product and the percentage as decoration. Graduate to trending the score once a single response can no longer visibly move it.
The metrics that lie
Average handle time as a target. As an observation it is fine; as a goal it rewards rushing. Every minute you shave off a conversation shows up somewhere else, usually as a reopen or a refund. Measure it if you're curious. Never bonus anyone on it.
Deflection rate on its own. It counts the visitor you helped and the visitor you exhausted as the same win. Deflection means something only next to what happened afterward: did they come back through email, or leave a review that answers the question for everyone else?
NPS on support interactions. "How likely are you to recommend us" after a shipping question measures the brand, the product, and the discount they just got. It is the wrong instrument for a single interaction. Keep it for the relationship, if you use it at all.
Raw ticket count. It punishes growth. More tickets in a month where sales doubled is good news wearing a scary hat, which is exactly why contact rate exists: it divides the scary number by the good one.
What AI answering first adds to the board
When an AI agent takes the first reply, a few new customer service metrics appear, and one old one changes shape.
AI resolution rate. The share of conversations the bot finished on its own. Count it only when the conversation ended without a handoff and without an abandonment. A visitor who closed the widget mid-answer is not a resolution, whatever the dashboard wants to call it.
Handoff rate. The share passed to a human. Read it as health, not failure. A bot with a zero handoff rate either has perfect knowledge or a blocked exit, and I have only ever seen the second. When handoffs climb, read those transcripts before you celebrate or panic. Sometimes it means harder questions are arriving; sometimes it means a knowledge gap; the transcripts will say which.
Knowledge gaps. Questions that found no good source in your content. This one barely qualifies as a metric. It is a to-do list wearing a chart, and each row is a page you should write or fix this week. In Hey Support this list is part of the insights view, and it is the report I would keep if I could keep only one.
Answer feedback. Thumbs on individual answers. The counts stay small, so skip the ratio and treat every thumbs-down as a lead: find the transcript, fix the source, same day. This is where an improve-answer loop earns its keep, because the correction becomes part of the knowledge and the same question stops producing the same bad answer.
The old number that changes shape is first response time. The bot answers instantly, so the blended average collapses toward zero and flatters the whole board. Measure the human reply time after handoff separately, because that is the number your escalated customers actually feel.
Instrumentation without a data team
You do not need a warehouse. Your support tool's built-in numbers plus one spreadsheet covers everything above. The spreadsheet gets a row per week and columns for the core four plus handoff rate. Filling it in takes ten minutes.
The other ten minutes are the ritual: one sentence on what moved and why, and one action for the week. Fix a source page, rewrite a bad answer, adjust a handoff rule. The weekly transcript read from our chatbot best practices pairs with this; the transcripts explain the numbers, and the numbers tell you which transcripts to read.
A concrete version, twenty minutes on a Monday: write the numbers down, answer three questions in one sentence each (what moved, why do we think so, what one thing changes this week), then open the three worst transcripts and check whether the story holds. The numbers suggest a cause; the transcripts confirm it or embarrass it. Skip the meeting; a two-person thread is enough.
The cheat sheet
| Metric | What it tells you | How it lies |
|---|---|---|
| First response time | Whether anyone is home | Averages bury the overnight gap; auto-replies game it |
| Resolution rate | Whether questions end answered | "Ended" is not "answered"; silence counts as success |
| Contact rate | Whether growth creates support work | Rarely; just keep the denominator consistent |
| CSAT | How the people who replied felt | Response bias; small samples swing on one bad day |
| Average handle time | Rough cost per conversation | As a target it rewards rushing and punishes care |
| Deflection rate | How many never reached a human | Counts the helped and the exhausted as the same win |
| NPS (after support) | Brand feeling, at the wrong moment | Measures the relationship, not the interaction |
| Raw ticket count | Workload | Punishes growth; needs a denominator to mean anything |
| AI resolution rate | What the bot finishes alone | Honest only if abandonment is subtracted |
| Handoff rate | Whether the seam to humans works | Read as failure, it tempts you to block the exit |
One more filter: match the board to your stage. In the first month after launch, ignore most of this and live in the transcripts and the gap list, because the numbers are too small to trend. Once volume settles, promote contact rate and handoff rate to the weekly review and demote the rest to occasional. The board should shrink over time as you learn which numbers you actually act on, and a shrinking board is a sign of health, not neglect.
Metrics are for deciding, not reporting. If a number cannot change what you do next week, stop collecting it. Start with contact rate and the weekly twenty minutes, and add the rest when a decision needs them. If you want to see where the board fits in a full rollout, the AI customer support guide walks the whole path.
Hey Support tracks resolutions, handoffs, and knowledge gaps for every bot, on flat pricing that never bills per answer, so a busy month on the board is only ever good news.



