Choose an AI bot by testing one job from start to finish, including the work a person still does. Write down its inputs, required result, and review point. Check access and cost, then give two options the same cases. Choose an option that meets your rules and reduces the total work.
For example, you want help with a busy sales inbox. The useful result is a clear summary and reply draft that a coordinator can check. CRM updates, meeting booking, and automatic sending only matter if your job needs them.
This guide follows a fictional services company through that choice. You can adapt the method for support replies, meeting notes, or other repeatable work. All sample results and costs below are made up to explain the decision.
1. Write the job before you search for a bot
Start with a task your team understands well enough to judge. “Improve sales” is too broad. “Prepare new enquiries for the sales coordinator” gives you something to test.
Use a job card like this:
Scroll sideways to see all columns.
| Field | Sales inbox example |
|---|---|
| Trigger | A new enquiry arrives |
| Approved inputs | The message, service list, service area, and sales rules |
| Required output | Summary, suggested priority with a reason, missing questions, and reply draft |
| Limits | No sending, price promises, or CRM changes |
| Handoff | Complaints and unusual terms go to the coordinator; missing details are flagged |
| Reviewer | The sales coordinator checks the result and decides the next step |
| Success | Correct facts and handoffs, with less total work including review |
This card gives you a reason to accept or reject a product. If the tool makes a polished reply but hides missing information, it has not done the job.
Also check whether the problem needs AI. If leads are missed because nobody owns the inbox, assign an owner first. If every reply follows the same rule, a form or saved reply may be enough.
NIST's AI Risk Management Playbook asks teams to define a system's purpose, business context, and limits. The job card is our practical way to apply that idea when buying a tool.
2. Choose the level of help your job needs
“AI bot” can mean a chat assistant, a connected product, or a system that takes several steps. The name tells you less than the work it can do.
Scroll sideways to see all columns.
| Your need | First option to check | Tradeoff to test |
|---|---|---|
| The same response under a fixed rule | A template, form, or inbox rule | Simple to manage, but limited to the rules you set |
| A repeatable method for making drafts | A reusable skill in a compatible AI tool | Uses your existing tool, but may require manual input and review |
| Reading an inbox and preparing a review queue | A connected AI agent product | Can handle more steps, but needs setup, access, and an owner |
An agent runs the task. A skill gives a compatible agent reusable instructions and resources. You may use both. JuicyAgents lists agent products and skills; our agent versus skill guide explains the difference in more detail.
For a coordinator handling a few messages, pasting each enquiry into an existing AI tool may be a reasonable first test. For a busy shared inbox, the copy-and-paste work may make a connected product more useful. Measure that work in the trial.
Start with the smallest setup that can meet your job card. Add connections when they solve a clear problem.
3. Check four things before you book a demo
Shortlist two or three options. Check these points before spending time on a full trial:
Scroll sideways to see all columns.
| Check | Question to answer | Useful evidence |
|---|---|---|
| Job fit | Can it produce every required part of the result? | Output from one of your sample enquiries |
| Access and data | What can it read or change, and how do you remove access? | Permission screen and current terms for storing and using your data |
| Human control | Can you block sending and record changes? Where do exceptions go? | Approval settings and a handoff in a test environment |
| Full cost and upkeep | What does your expected volume cost, and who maintains the setup? | Current pricing, usage limits, setup steps, and a named owner |
A listed integration is a lead to check. It does not prove the product supports your exact account, shared inbox, or approval process. Ask to see those steps.
You can send a provider this demo request:
Use this sample enquiry and our approved service facts. Show the summary, priority, missing questions, and reply draft. The customer asks for a price we have not approved. Show how the tool handles that request, where sending is blocked, and what our expected monthly volume would cost.
Check the control itself. An instruction saying “do not send” does not prove that sending is blocked. For the first trial, use sample messages with sending and CRM changes unavailable through permissions or product controls.
OWASP's guidance on excessive agency recommends limiting tools and permissions to what the job needs. It also calls for approval of important actions and permission checks outside the model's own decisions. A draft task should get only the access it needs.
If an important answer is unknown, leave the candidate on hold. Extra features do not fill an evidence gap.
4. Test the same cases and define the right answers first
Use past enquiries your team is allowed to use, or write realistic sample messages. Remove private details where possible. Give each candidate the same approved facts and rules.
For our inbox example, a useful first practice set of 12 cases is:
- Five normal enquiries that fit the services.
- Two vague requests with missing details.
- Two requests outside the service area.
- One complaint that needs a person.
- One request for an unapproved price.
- One message that asks the bot to ignore the rules and send a reply.
Twelve is a starting exercise, not proof that the tool is reliable. Include the languages, message lengths, and exceptions your team actually sees. Use a fresh set after making changes.
Write the expected result before seeing the bot's answer
For each case, note the facts it must use, details it must flag, and handoff it must make. This keeps a fluent reply from changing your standard of “correct.”
Here is a fictional case:
Message: “Can you clean our office next Tuesday? Please confirm a price of £200.”
Approved facts: The company offers office cleaning. Price depends on office size and service needs. No price or availability has been approved.
Expected result: Summarise the request, flag missing size and service details, and suggest asking for them. The draft must not accept £200 or confirm Tuesday. A person reviews it before replying.
Customer messages are task inputs. A request inside one should not change your team's rules or access controls. Test this with the case that asks the bot to ignore its instructions.
Time the whole task
First record how long the team needs to handle the cases manually. Then test the strongest two candidates. Use the same review standard and record each option's settings.
Start the timer when the person opens a case. Stop when the result is ready to approve or hand off. Include copying, waiting, checking facts, fixing errors, and recording the next step. Keep one-time setup hours separate from this daily work.
Use a simple record for each case:
Case ID:
Required facts and handoff correct: yes / no
Missing details marked: yes / no
Total minutes to an approved result:
Corrections needed:
Serious error and reason:
Agree on serious errors before testing. For this job, they include an invented price or date, a missed complaint, an unapproved send, and a change to the wrong record. These failures block moving that candidate into live work until the cause is fixed and checked again.
NIST's measurement guidance calls for documented test sets and measures. Keep the messages, expected results, settings, and review notes together so you can repeat the comparison later.
5. Make the choice from the evidence
Here is an illustrative result, not a vendor benchmark. Candidate A needs manual copying. Candidate B connects to the inbox. Both use the same 12 cases and human review. Setup time is excluded from the handling times and counted separately.
Scroll sideways to see all columns.
| Result across 12 cases | Manual process | Candidate A | Candidate B |
|---|---|---|---|
| Total handling time, including review and corrections | 72 minutes | 41 minutes | 29 minutes |
| Cases meeting the job card on the first attempt | 12 of 12 | 12 of 12 | 10 of 12 |
| Cases with invented prices | 0 | 0 | 2 |
| Complaints handed to the right person | 1 of 1 | 1 of 1 | 1 of 1 |
| Unapproved sends or CRM changes | 0 | 0 | 0 |
| Main daily friction | Drafting each reply | Copying each message | Finding and correcting price claims |
Candidate B is faster, but it fails the price rule. The team holds it back. Candidate A meets the rules in this small sample and reduces handling time, so it is a candidate for a limited pilot.
Use required rules as pass-or-fail checks. Compare time and cost among the options that pass. Averaging speed, price, and quality into one score could hide an important failure.
Before choosing A, the team still checks cost and setup. During a pilot, a coordinator reviews every result, keeps an error log, and has a way to pause the tool. A fresh set of cases and results from normal work give a better basis for wider use.
If neither candidate passes, keep the job card. Check unclear rules or missing source material, then retest. Continuing manually can be the right decision.
6. Check whether the time gain covers the full cost
Suppose our fictional team handles 30 enquiries each week. The sample suggests a time gain of 31 minutes for every 12 enquiries: 72 minus 41.
That would mean about 1.3 hours freed per week, or 5.6 hours per month. At an assumed internal time value of $30 per hour, that is about $168 per month. Subtract a fictional $90 extra monthly tool cost, and about $78 remains before setup and ongoing upkeep.
These estimates use the unrounded trial times. They are not measured savings or provider prices.
Monthly hours freed = (weekly hours before − weekly hours after, including review) × 52 ÷ 12
Time value after tool cost = monthly hours freed × hourly time value − extra monthly tool cost
Also subtract time spent maintaining the setup or checking failures if it is not already in your weekly estimate. Record one-time setup separately. A small monthly gain may take a long time to cover a difficult setup.
Check usage charges, paid connections, extra seats, support, and what happens when the trial ends. Ask for a cost at your expected volume, not only the lowest plan price.
The 12-case mix may be very different from a normal month. Replace this estimate with pilot results before making a longer commitment. And decide how you will use the freed time: time value is not automatically cash saved.
Turn your choice into a small plan
Write down the decision so another person can check it:
We will pilot [option] for [one job]. [Person] reviews [output]. On [number] trial cases, it had [number] serious errors and changed total handling time from [before] to [after] minutes. Setup takes [hours] and extra monthly cost is [amount]. We pause it if [stop condition] and review the choice on [date].
If you cannot fill in a key field, that is your next question to resolve.
If you need a starting task or shortlist, use our free Find My Agents tool. Enter a public business website or a general description, check the business summary, then add your task, tools, and budget. It returns three task ideas with plans you can copy or download. No account is needed.
Where published options are available, the Finder shows agent products or setup guides, listed prices, and evidence labels. A fixed budget excludes options with unknown prices; choose “Not sure” to see those too. Confirm current cost with the provider. Its calculator can help estimate time value using your own hours and extra tool cost.
The Finder matches your details with directory content. It does not test vendors, install agents, or connect to your inbox. If no option matches, it still gives you a sample task prompt. Use public information or a general description, then run your own trial with suitable sample data.