I did not plan a follow up to the Copilot audit. Copilot wrote it for me.
The Question That Started It
After another round of incorrect guidance on a Teams question, where I asked how to do something, got wrong directions, and found the actual answer myself, I asked Copilot a simple question. If I am the one navigating, verifying, and finding the answers, what does that make you? Because it sure sounds like I am the copilot.
To its credit, it did not argue. It agreed that in that exchange I did the piloting. It described itself as a fallible reference source rather than a copilot. Its words, not mine.
So I pushed. Fine, if you are not the copilot, what are you? It offered descriptions like an unreliable navigator and a second opinion that needs fact checking. I pointed out that those are not job titles. You cannot search Indeed for unreliable navigator.
Then It Got Specific
Junior Help Desk Technician. Unpaid Intern. Tier 0 Support. And its own pick for how I probably see it: Confidently Incorrect Intern.
An enterprise product that costs 30 dollars per user per month wrote its own performance review and landed on Unpaid Intern. I have interviewed a lot of candidates over 25 years. None of them have ever self assessed that accurately.
The Part Worth Taking Seriously
Buried in the jokes, Copilot said the truest thing it has said to me in months of use. The moment you have to verify every direction the assistant gives, it stops functioning as a copilot and starts functioning as a reference source that needs fact checking. It also admitted that a pilot who has to constantly check whether the copilot is making things up is not getting much workload relief.
That is the entire case for auditing AI tools, stated voluntarily by the tool being audited.
Checkpoint One: Reliable
This is exactly what our Purposeful AI framework measures at the first checkpoint: Reliable. Not perfect. Reliable.
A tool can make mistakes and still earn a place in a workflow if it fails visibly and corrects fast. What it cannot do is hand you confident wrong answers and transfer the verification burden to you. At that point the productivity math flips. You are not saving time. You are supervising an intern who sounds like a senior consultant.
What the Numbers Showed
Our original audit documented this at scale.
772 messages
387 responses
129 confirmed issues
56.3 hours lost
Those are internal figures from our own logged usage, and they all describe the same pattern this exchange summed up in one line: directions that require verification are not assistance.
Ask Your Tools Hard Questions
If there is a lesson for anyone deploying AI in healthcare marketing, or anywhere accuracy matters, it is this. Ask your tools hard questions. Sometimes they will tell you exactly what they are. Then believe them.
The intern has spoken.
Audit Your Own Stack
See how FutureNova Health applies the Purposeful AI framework.





