WM Blog · Clara

Duplicate Rules Starve Your AI Forecasts

Your revenue ops lead keeps entity-matching thresholds frozen after the first MDM rollout. The result is AI models that treat near-identical customers as separate accounts and produce forecasts that miss by double digits.

Fragmented duplicate records disrupting an AI forecast graph on a dark background

Revenue operations teams still treat master data as a one-time project with static matching rules. Once the initial deduplication pass finishes, thresholds for name, address and ABN similarity never move again.

Six months later the AI demand model sees two versions of the same mining customer with different contact histories. It predicts separate churn risks and recommends contradictory pricing actions.

The finance controller then questions why pipeline accuracy dropped after the new forecasting tool went live. The root cause sits in the CRM sync layer where a 0.85 similarity cutoff silently splits records.

Procurement data suffers the same split. A supplier appears under three legal entities because invoice addresses differ by a single digit. The AI spend model therefore shows phantom concentration risk and flags the wrong vendor for renegotiation.

No one owns the ongoing calibration. The data steward reports to IT architecture, not the P&L owner who feels the forecast error. Quarterly cleans only catch the obvious exact matches and leave the grey-area duplicates untouched.

Competitors in Perth and Brisbane run nightly entity-resolution jobs that feed live scores back into their models. Their forecasts tighten while yours drift because your matching logic froze at go-live.

Fix the mechanism. Tie the revenue ops lead’s bonus to measured duplicate rates inside active AI training sets, not to the number of records declared golden. Re-run matching daily with feedback from model error logs and adjust thresholds before the next forecast cycle closes.

Data AI Operations Governance