What a credible process considers
These elements keep the analysis tied to live use.
- eligible test population
- random assignment
- policy compliance
- outcome maturity
- win rate
- revenue per appointment

Design a practical A/B test for machine-learning-driven sales assignment using eligible populations, mature outcomes, adoption, and guardrails.
Start with historical appointments, realistic eligibility rules, and outcomes that have had time to mature.
Randomly assign eligible appointments to the current policy or the model-guided policy, preserve the same operating constraints, wait for outcomes to mature, and compare win rate and revenue per appointment. Track whether recommendations were actually followed.
The useful prediction is not merely whether an opportunity will close. The system must estimate how the expected outcome changes across the eligible representative choices, using only context available at assignment time.
The ranking becomes an operational policy only after eligibility, capacity, territory, and workflow requirements are applied.
These elements keep the analysis tied to live use.
These shortcuts can make model performance look better than it is.
| Measurement | Why it matters | How to use it |
|---|---|---|
| Held-out log loss | Tests whether the method is accurate, stable, calibrated, and economically useful. | Compare appointments that have had enough time to reach an outcome |
| Brier score | Tests whether the method is accurate, stable, calibrated, and economically useful. | Compare similar periods, lead sources, and assignment approaches |
| Precision-recall performance | Tests whether the method is accurate, stable, calibrated, and economically useful. | Compare appointments assigned to recommended representatives |
| Recommendation alignment | Tests whether the method is accurate, stable, calibrated, and economically useful. | Review prediction accuracy alongside actual sales and revenue |
| Revenue per appointment | Tests whether the method is accurate, stable, calibrated, and economically useful. | Check for changes in lead mix and team performance before changing your approach |
Isotope Labs organizes your appointments, representative assignments, and outcomes into a history we can evaluate, using information available before each assignment.
We test predictions against historical outcomes kept separate from model training to identify reliable matches between opportunities and sales representatives.
We evaluate historical assignment scenarios that reflect representative eligibility, territories, availability, and workload limits.
See which recommendations your team uses and how those appointments perform, including close rate and revenue per appointment as outcomes become available.
Randomly assign eligible appointments to the current policy or the model-guided policy, preserve the same operating constraints, wait for outcomes to mature, and compare win rate and revenue per appointment. Track whether recommendations were actually followed.
Use information known before assignment, a stable representative identifier, appointment date, and a mature won or lost outcome. Use reliable information that was available before the appointment was assigned.
Use held-out data, compare against simple baselines, review ranking and probability metrics, and simulate the actual assignment constraints.
Monitor recommendation coverage, alignment, mature win rate, revenue per appointment, calibration, agent capacity, and changes in source or opportunity mix.
Use capacity-constrained sales assignment to rank representatives while respecting appointment limits, schedules, territories, and service levels.
Strategy guideMeasure revenue per appointment alongside win rate, contract value, maturity, source mix, and recommendation alignment.
Strategy guideReview the fields, definitions, and history needed for a credible evaluation.
Share a representative CRM export or connect your data. Isotope Labs will evaluate data readiness, representative-performance variation, and the potential value of model-guided assignment before a live rollout.