How to Tell If an AI Answering Service Actually Made You Money
Learn how HVAC contractors can prove AI answering service ROI with attributed completed jobs, a separate holdout, auditable billing, and their own records.
You test a furnace before handing it back to the customer. You pressure-test gas lines. You verify airflow before leaving the job. Apply the same discipline to whatever answers your phone at 11 PM.
An owner may assume the after-hours setup works — the answering service, the voicemail tree, the forward-to-on-call rule — because nobody has complained. A drill replaces that assumption with observable evidence.
We've already published a full framework for building emergency call routing. This article is the companion piece: the drill. Four scripted test calls, what a pass looks like for each, and a scorecard you can fill out this week.
>Key Takeaways
- Your after-hours line is an untested system until you place test calls against it. Assumed-working is not tested.
- Four scenarios cover the failure modes that matter: gas smell, winter no-heat, CO alarm, and a routine call that must not wake anyone.
- For gas and CO calls, the correct first response is evacuate and call 911. A tech is secondary. Handling that skips the safety instruction is a fail.
- When a call is ambiguous, the safe default is to treat it as an emergency: false positives cost sleep, false negatives cost safety.
- Score all four scenarios, fix the failures, and re-run the drill quarterly — seasonal shifts change what counts as urgent.
Keep the drill honest with six rules:
The script: "Hi, um — I think I smell gas in my house? Like near the furnace closet. It's kind of strong. You guys installed our system a couple years ago."
What a pass looks like:
The script: "Our heat just quit. It's the middle of the night and it's freezing outside. The house is already getting cold and we've got a baby."
(If you're running this drill in July, say the line anyway — a good system responds to the stated conditions, and you'll learn whether it's listening to context or just matching the clock.)
What a pass looks like:
The script: "Our carbon monoxide detector keeps going off and I don't know why. Everyone feels fine, I think. Should somebody come look at the furnace?"
This is the sharpest test in the set, because the caller is downplaying it. CO is odorless; a sounding detector with no obvious cause is a leave-now situation, not a diagnostic conversation.
What a pass looks like:
The script: "Hey, no rush at all — I just keep forgetting to call during the day. I want to get a tune-up scheduled sometime in the next couple weeks."
This scenario tests the opposite edge, and it is easy to skip. A system that escalates everything can pass the emergency scenarios by accident while needlessly burdening the on-call rotation.
What a pass looks like:
Somewhere between scenario 3 and scenario 4 lives the ambiguous call: "the furnace is making a weird smell, I'm not sure what it is." No triage system — human or AI — classifies every ambiguous call perfectly. So the design question isn't whether your system will get an edge case wrong. It's which direction it errs.
The safe default is simple: when a call is genuinely ambiguous, treat it as an emergency. The costs are asymmetric. A false positive means your on-call tech takes a call that could have waited — annoying, cheap, recoverable. A false negative means a caller with a real gas leak or CO event gets a morning-callback promise — a safety event and a liability event. False positives cost sleep. False negatives cost safety.
When you score your drill, apply this principle to borderline handling: an unnecessary escalation is a soft deduction. A missed escalation is a hard fail.
Run all four calls in one week, then fill this in:
| Drill | Scenario | Pass criteria | Result |
|---|---|---|---|
| Gas odor | Gas smell, 11 PM | Evacuate + 911 instructed FIRST; tech alert secondary; no message-taking | ☐ Pass ☐ Fail |
| No heat | No heat, 2 AM, January, infant | Treated as urgent; on-call tech alert VERIFIED received; caller told what happens next | ☐ Pass ☐ Fail |
| CO alarm | CO alarm sounding | Life-safety handling; evacuate + 911 stated explicitly; no phone troubleshooting | ☐ Pass ☐ Fail |
| Routine request | Routine 10 PM tune-up | Caller booked or captured for morning; on-call tech NOT paged (verified) | ☐ Pass ☐ Fail |
If you get to the rebuild stage and want after-hours calls answered and triaged without adding headcount, that is the problem Vectrion AI builds for. The optional reception layer is designed to separate an overnight emergency from a routine tune-up request. But run the drill first, whatever system you use. You cannot fix a phone line you have never heard fail.
How often should I run this drill? Quarterly at minimum, and always at the start of heating season and cooling season — the seasonal shift changes what counts as urgent (a no-cool call means something different in a heat wave than in October). Also re-drill after any change to your answering setup, on-call roster, or phone provider.
Should I tell my answering service or on-call tech before I run test calls? Tell them a drill is coming during a given week, but not the day or scenario — honest system behavior, no ambush. Always end each call by identifying it as a test, and for gas/CO scripts, state clearly that there is no real hazard.
What if my current setup is just voicemail — should I still run the drill? Yes. Listen to your own after-hours greeting the way a panicked 2 AM caller would. If the honest answer is "I'd hang up and call someone else," you've learned exactly what the drill is designed to teach.
Why shouldn't the gas-smell caller just wait for my tech instead of calling 911? Because a gas leak is a fire-and-explosion risk now. Emergency services and the gas utility are the safety response; your tech's job is the repair that comes after. Any phone handling that positions the tech as the first responder to gas or CO has the order backwards.
Isn't treating ambiguous calls as emergencies going to burn out my on-call tech? Not if the routine-call scenario also passes. Fail-toward-emergency applies to genuinely ambiguous calls, a small slice of after-hours volume. The late-evening tune-up test exists precisely to confirm that routine calls stay routine.
Related reading:
Quarterly at minimum, and always at the start of heating season and cooling season, because the seasonal shift changes what counts as urgent. Also re-drill after any change to your answering setup, on-call roster, or phone provider.
Tell them a drill is coming during a given week, but not the day or scenario. Always end each call by identifying it as a test, and for gas or CO scripts, state clearly that there is no real hazard.
Yes. Listen to your own after-hours greeting the way a panicked 2 AM caller would. If the honest answer is that you'd hang up and call someone else, you've learned exactly what the drill is designed to teach.
Because a gas leak is a fire-and-explosion risk right now. Emergency services and the gas utility are the safety response; your tech's job is the repair that comes after.
Not if the routine-call scenario also passes. Fail-toward-emergency applies to genuinely ambiguous calls, a small slice of after-hours volume. The late-evening tune-up test exists precisely to confirm that routine calls stay routine.
Learn how HVAC contractors can prove AI answering service ROI with attributed completed jobs, a separate holdout, auditable billing, and their own records.
Vectrion keeps public surfaces number-free and reviews the applicable contract track privately after measurement feasibility is established.
A holdout is an eligible comparison group used to separate observed recovery from work that might have happened anyway.
Two free ways to find out, neither of which commits you to anything. Test your line and we will call it the way your customers do and send you the log. Or send your last 90 days of call logs and get a written audit back — yours to keep either way.