CoverTurn.
Ended deliberately AI product 2026

An autonomous voice agent.

Built end to end and run against 582 real businesses. Instrumented before launch rather than after, which is the only reason it taught me anything. The telemetry disproved my own hypothesis, so I ended the programme instead of defending the build.

Scope Telephony + API integration + Data pipeline + Evaluation Outcome Ended on the data
582
businesses called, autonomously
374
answered and held a conversation
1.45s
median end to end latency
13¢
cost per call, logged from day one
01 · The build

A real system, not a demo on a laptop.

The agent placed live calls over the public telephone network, held a real conversation, and booked meetings straight into a calendar over API while still on the line. Behind it sat a webhook pipeline that extracted seven structured fields from every single call into a database: who answered, what they said, the objection, the outcome.

All of it built solo. Telephony layer, booking integration, extraction pipeline, and the conversation design itself, iterated across eight versions against logged transcripts rather than against my own opinion.

02 · The decision that mattered

Instrument before launch, not after.

Most people bolt analytics on once something underperforms. By then you are reconstructing a story from fragments. So latency and cost per call were logged from the first call, alongside the structured outcome of every conversation.

That decision is the entire reason this project produced knowledge instead of a bill.

03 · Being wrong, with evidence

The data contradicted me, so I followed the data.

I held the belief everyone in voice AI holds: latency causes detection, detection causes hangups. I believed it enough to spend weeks on it, pulling median response time down across eight successive versions.

Then I split the 374 answered calls by latency and checked. Conversations on the fast calls ran a median of 23 seconds. Conversations on the slow calls ran 27. The slower ones lasted longer. My hypothesis was wrong, and I had spent weeks acting on it.

What the data actually showed was a cliff between ten and fifteen seconds, where conversation survival fell from 89% to 56%. A third of every conversation died before the pitch even began, which is an opener problem, not a model problem. And 102 of the 374 answered calls ended in silence rather than a hangup: people did not reject the agent, they disengaged and it timed out. That is conversational design.

04 · The outcome

Knowing when to stop is part of the job.

I ended the programme. The constraint was segment fit, and no amount of prompt engineering fixes a segment that does not take phone calls. Continuing would have meant optimising a system that could not win.

Total cost of establishing that: $90.

The transferable version, and the reason this sits on the site alongside the things that worked: define what success means before go-live, instrument for it from day one rather than retrofitting, and be willing to give an honest reading of your own numbers when they disagree with you.

A deployment that runs but is not used has not succeeded.

Scope
Telephony + API + Data + Evaluation
Year
2026
Versions shipped
8, against logged transcripts
Status
Ended deliberately, on the data
Next case

Check the Lease.

Read the next case
Get in touch
Let's build something worth measuring.