The first ROI question
How we measure whether an automation is actually working, before a client has to ask.
· 5 min read

Every AI vendor has a good demo. Fewer have a system still running six months later.
We've sat in on enough post-mortems to see the pattern repeat: a pilot gets approved, a small team builds something impressive in a sandbox, leadership sees a great presentation — and then nothing changes in how the business actually operates. The project quietly stops getting mentioned in stand-ups. Eventually someone asks "whatever happened to that AI thing?" and nobody has a clean answer.
The technology was rarely the problem. The gap was almost always somewhere else.
The demo environment lies to you
A pilot usually runs on clean, hand-picked data. Production runs on whatever your business actually generates — inconsistent formats, missing fields, edge cases nobody thought to test. A system that performs beautifully on fifty curated examples can fall apart on the five hundred messy ones it meets in week one..
This is why we test every system against a client's actual historical data before it goes anywhere near production. If it can't handle the real mess, it isn't ready — no matter how good it looked in the pitch.
Nobody owns it after the handoff
Pilots often get built, presented, and then handed to a team that had no part in building it. Without context on why decisions were made, that team has no way to maintain, debug, or trust the system. It gets used cautiously, then rarely, then not at all.
We stay engaged through the first month a system is live — not to keep billing hours, but because that's when the real edge cases show up, and someone needs to be responsible for fixing them.
There was never a clear definition of done
Build us an AI agent is not a scope. Without a specific operational target a metric, a workflow, a task that gets fully handled a project can run indefinitely without ever being finished, because there was nothing concrete to finish.
Before we write a line of code, we agree on exactly what the system needs to do, on what data, with what outcome. That's what done means for us — not a working demo, but a defined task fully handled in production.
What we changed about how we work
Every engagement now starts with real data, not sample data. Every system has a named internal owner on the client side before we build anything. And nothing gets marked complete until it's running unattended in production for at least two weeks.
It's a slower start than most agencies promise, which is also why our clients are still successfully running what we built a year later, instead of asking what happened to it and why it works so well.


