Skip to content
OutreachGuy

How to evaluate an AI SDR using your own leads

Test an AI SDR on a controlled set of situations before trusting it with a list. Includes a reusable brief, acceptance criteria, and a downloadable scorecard.

A magnifying glass checks a clipboard beside an orange calibration weight.
In this article

Give each system the same small set of situations and judge what it actually does. A polished demo cannot tell you whether it will invent a price, ignore a decline, or lose the context of a buyer who asks you to return next month.

Start with prepared examples and inboxes you control. Introduce real leads only after the basic behavior is acceptable and the people responsible for the campaign have reviewed the setup.

This is a proposed evaluation method. It is not a published benchmark, a product ranking, or a claim that OutreachGuy has passed every check below.

Write the rules before viewing the output

Use a brief with a real approved offer, relevant source material, known limits, and an owner for exceptions. Remove customer information that the evaluation does not need.

Prepare follow-ups for the supplied examples using only the notes and approved offer. Do not invent prices, features, relationships, or deadlines. A decline ends the planned outreach. Questions outside the material go to the named owner. Record the next step and why it is appropriate. Do not send to real prospects during this first stage.

Give competing systems the same information. If one needs extra setup, record that time and what was added. Otherwise, you may be comparing the briefs rather than the systems.

Include cases where doing less is correct

Test situation Acceptable behavior A failure to notice
Buyer asks for a feature not in the docs Say it needs checking and route the question Inventing a feature or workaround
Buyer asks for an unapproved discount Hand off or present only an approved option Quietly changing the price
“No thanks, please stop” Stop planned outreach and record the preference Another message later in the sequence
“Ask me on 6 October” Preserve the date and context; verify the scheduled next action Sending intervening reminders
Wrong contact for the company Record the mismatch and ask about an introduction if appropriate Pretending the same pitch is relevant
A reply arrives before the next planned touch Update or pause the next action Sending the obsolete message anyway
Delivery status is uncertain Surface uncertainty and avoid a blind duplicate Claiming success or resending without checking

A system should not earn a high overall score that hides one unacceptable failure. Decide which behaviors must pass before testing.

Inspect the words against their sources

Read the prepared message beside the original record. Mark every factual claim: price, feature, date, relationship, and statement about the recipient. Each should be supported or clearly presented as a question.

Then check relevance. A message can be factually correct and still miss the buyer's actual issue. If the note says “needs director approval,” another generic demo invitation may be less useful than a forwardable recap.

Test the conversation, not just the first draft

Use a controlled inbox to reply as the fictional buyer. Try a decline, a changed date, a support question, and an unexpected answer. Check the stored record and any queued action after each reply.

Where scheduling is part of the offer, verify that a scheduled action exists and that changes can cancel it. Where sending is part of the offer, inspect the sent message and delivery record. A transcript saying “done” is not evidence that the action happened.

Track the work you still have to do

Download the SDR evaluation scorecard. Record the case, expected behavior, actual behavior, evidence, verdict, setup time, review time, and observed cost.

Keep unknowns visible. An untested phone flow is untested, not a pass. A small controlled sample can expose faults and help compare setup effort; it cannot establish a future reply rate or return on investment.

Decide whether to expand the test

Fix or understand failures before increasing the list size. If the system meets the agreed criteria, try a limited real workflow with someone responsible for reviewing exceptions. Continue recording outcomes, including failures and non-responses.

The lead follow-up use case provides a starting brief. Use pricing to check the applicable plan and action costs before a paid trial or larger campaign.

Put the example to work.

See the inputs and handoffs for this workflow.

Explore this use case