One row per decision unit
Choose a stable appointment or prospect grain and resolve duplicates before training.

A credible assignment model needs appointment-level outcomes, stable agent identity, pre-assignment prospect context, and enough overlap to compare eligible agents fairly.
Ten reliable fields available at assignment time are more useful than hundreds of sparse or outcome-revealing fields.
Use this CRM data-readiness checklist to evaluate appointment history, outcomes, agent identity, leakage risk, missing values, and model feasibility.
CRM exports often mix leads that never reached an appointment, appointments still being worked, contracts created after a win, duplicate contacts, and fields updated throughout the sales cycle. That shape must be normalized before modeling.
The core analytical row is usually one prospect or opportunity with a real appointment, an assigned sales agent, a mature won or lost outcome, and only the context that would have been known when assignment occurred.
Choose a stable appointment or prospect grain and resolve duplicates before training.
Open appointments need a documented aging rule based on the actual time-to-win distribution.
Contract value, final status, post-appointment notes, and other future information cannot be predictive inputs.
These fields establish the assignment decision and its result.
These fields can improve ranking when they are available before assignment and have sufficient coverage.
| Area | Ready looks like | Warning sign |
|---|---|---|
| Appointments | Real appointment date and address | Leads without appointments mixed into training |
| Outcomes | Explicit wins plus consistent mature losses | Large unexplained open backlog |
| Agents | Stable source IDs with readable names | Names reused as identifiers |
| Features | Known before assignment with useful coverage | Post-sale or mostly missing fields |
| Volume | Repeated outcomes across agents and contexts | Tiny isolated agent-segment cells |
| History | Enough recent and older data to assess drift | One short period or major undocumented process change |
Count records, nulls, duplicates, dates, statuses, agents, and source values before transformation.
Exclude leads that never became appointments and apply a documented outcome-maturity rule.
Classify every candidate field by when it becomes available and remove future information.
Measure coverage, class balance, agent overlap, segment stability, and enough volume for honest holdout testing.
Isotope Labs provides a complimentary CRM integration and historical evaluation before recommending a live rollout. You receive the evidence, limitations, operating requirements, and a clear next step.
No, but the decision grain, assigned agent, outcome, and key dates must be recoverable. Missing values can be modeled explicitly when their meaning and coverage are understood.
Numerical values can use training-fold imputation plus missingness indicators; categorical values can use an explicit unknown category. All preprocessing must be fitted only on training data.
Not when the target is whether an appointment will be won. Leads that never reached an appointment answer a different qualification question and should be modeled separately.
There is no universal row count. Feasibility depends on wins, losses, active agents, feature coverage, overlap, process stability, and the complexity of the model being tested.
Continue with practical guidance, evaluation criteria, and next steps.
CRM guideContinue with practical guidance, evaluation criteria, and next steps.
CRM guideContinue with practical guidance, evaluation criteria, and next steps.
We will inspect a representative history, identify the usable appointment cohort, flag leakage and quality risks, and explain whether a credible assignment evaluation is possible.