Ways to Build Trust in Your Agents
TL;DR;
- Getting out of the loop is what frees up your time - The more work an agent does without you watching, the more time you get back.
- But you have to trust it first - Nobody hands a task over on faith, and trust in an agent is earned rather than assumed.
- Three ways to build that trust - Work in the chat, automate the task and review the output, or prove it with an eval suite.
I recently posted about how I rebuilt my blogging and social media pipeline so that I could involve AI agents as assistants while staying in full control of what is actually published. One of the key themes people picked up on was that automating your process gives a practical way to build trust in your agents over time by gradually reducing the amount of oversight and review you need to give them.
This got me thinking about other ways to build trust in AI agents, and I wanted to share my ideas with you.
The basic idea is that trust in an AI agent gets built in levels. Each level takes more effort than the one below it and gives you more confidence in return.
Level 1: Work in the chat
Open a chat, write a prompt, and push on it. Correct what comes back, run it again, adjust the wording. When you’re consistently getting what you wanted, save the instructions as a skill so you can call it up instead of retyping it.
This is known as working “in the loop”, or in-loop for short. It means you are at the keyboard guiding the AI at every step.
The whole value here is speed of feedback. You change one line and see the effect immediately, with nothing to rebuild in between. That’s also the limit: you’re the one starting it every time, and the results are only as repeatable as your memory of what you typed last week.
Use this when you’re still working out what you want. Most ideas should start here and plenty should stop here.
Level 2: Automate it, and review the output
Put the task on a trigger or a schedule, and read the output before it goes anywhere. The computer does the tedious part. You approve. Over time, as the agent keeps getting it right, you loosen the review: check every one, then check one in five, then check the exceptions.
This is known as working out-loop. Instead of being in the loop, the AI works on it’s own and you only step in as a reviewer. Getting out of the loop is the key to unlocking the potential of AI to free up your time. The more work you can trust the AI to do for you, the more time you have to work on other things.
Two costs. There’s setup work, though less than people expect. The bigger one is that your feedback loop slows right down. A daily task gives you one look per day, so a change you make on Monday isn’t confirmed until Tuesday, and three changes deep you’ve lost a week.
Use this when you know the task and you know what good looks like, but you’re not confident the agent has understood it yet. The review is how you find out, and it’s also what makes it safe to find out.
Level 3: Examples and an eval suite
Collect a set of real examples, with the output you’d want for each. Then build a way to run your prompt or skill against all of them and check the results. That’s an eval suite.
An AI eval suite (short for evaluation suite) is the AI development equivalent of a traditional software test suite. Because generative AI and LLMs are non-deterministic—meaning they can give different, open-ended answers to the same prompt—you cannot test them with standard pass/fail unit tests. An eval suite provides a structured, repeatable framework to measure the quality, accuracy, safety, and performance of an AI system. (from https://amplitude.com/explore/analytics/what-is-an-ai-evaluation)
This is a more advanced (but under-utilised) use of AI and you should expect it to be time consuming. Assembling good examples takes time, deciding what “correct” means for each one takes longer, and eval work has a habit of turning into an interesting project of its own that never quite gets back to the original task.
The payoff is that you can change things quickly and know nothing broke while you weren’t looking. Rewrite half the prompt, run the suite, see exactly which cases moved. Same when a new model comes out and you want to know whether switching costs you anything.
Use this for high-leverage cases where the prompt/skill will be used again and again in lots of different situations, highly critical places like parts of a website signup process you need to be sure, or when you expect to change the underlying model and want evidence rather than a feeling.
Picking a level
Which level to start at can be tricky at first but as you gain experience it will get easier to decide which strategy works for you. A common approach is to climb as far as you need to and stop there; work an idea out with an agent at level 1, then schedule it to run regularly once the big problems are sorted.
Level 2 is often enough on its own: the task runs, you read the output, confidence builds, and the eval suite never becomes necessary.