Thirty-seven minutes of a ServiceNow panel session, transcribed live, speakers named as they spoke, first words back in under six seconds. Getting there took two failed attempts and a fight with a spoon.
Last time I told you about the second quantum leap; scratch my own itch, build on top of my own experience. That path led me into constructing a completely fake client, hosting a pretend virtual workshop with a room full of people who don't exist. I know, Alice in Wonderland levels of bizarre.
Before I dive into the gory detail, let me explain how I got here.
Pretty much everything in delivery starts as a conversation; so the first test was transcription. And for ServiceNow discovery work the transcript is where everything starts; capture the conversation right and the service models, workbooks and decisions downstream all have a foundation.
For something realistic to transcribe I grabbed a public YouTube video, a recorded panel session, no client anywhere near it.
Question 1. Can a locally hosted AI accurately transcribe? Yes to this question would be a big win; local means completely private.
Answer; Yes. Local Whisper chewed through the audio at eleven times realtime, accurate enough to work with. BUT (and it's a big but) the setup process was like fighting your way out of an escape room blindfolded with nothing but a spoon. Just not practical.
Question 2. Can it tell me who said what? Worse again. Speaker identification (I learned it's called diarisation) locally meant a Jenga of Python dependencies, version pinning, and models you have to request access to; an evening of fighting it produced something that worked, but only as a batch job after the meeting ended and even then, not very accurately. OK for research. Useless in a live room.
Question 3. Could a commercial engine do it live? This one was a bit of a shocker. Thirty-seven minutes of that video transcribed at true realtime, speakers labelled as they spoke, first words back in under six seconds. Speaker recognition wasn't completely accurate all the time but it was good enough and a follow-up speaker identity review improved the results significantly. That was the moment this stopped being theoretical.
Confession time, it wasn't quite as straightforward as I'm making out. Poor Claude and Codex sweated for hours comparing the cost/benefit/performance of assorted commercial speech to text providers, testing a bunch of them while I sat back and watched before we settled on the right one.
Which brings us back to the fake client, because the next test needed something to listen to. Real workshops were never an option; you don't run experiments on clients. To be clear, the day this is proven and a client invites it in, I'll be in there like a rat up a drainpipe. But that's the point; proven first, invited first. So I wrote a workshop from scratch; a fully scripted discovery session for an application that doesn't exist.
How do you invent a roomful of people? Next time; casting my imaginary friends.
Have you fought the local-versus-cloud battle yourself? Tell me how it went in the comments.
First published on LinkedIn, 13 July 2026. The conversation lives there; the writing lives here.