An internal customer-health scorecard was supposed to predict renewals. Instead, teams manually re-verified its data before every client meeting — and some had stopped sharing its reports entirely. I led the research and design workstream that exposed why, and validated the redesign that shipped.
Inside one of the world's largest CRM platforms, a customer-health scorecard aggregates each account's product adoption, customer expertise, and technical health into a single score. Customer Success Managers, Account Executives, and Renewal Managers all lean on it — to prepare business reviews, forecast attrition risk, spot expansion opportunities, and build renewal strategy 90–120 days out.
The stakes are high: the score feeds forecasting, pricing conversations, and — for some roles — compensation. But by the time we started, the tool's core asset was gone. Not its data. Its credibility.
Scoring gaps — on-premise deployments and newer AI product SKUs that the model simply didn't track — produced false "0" scores on healthy accounts. Users couldn't tell a genuine adoption gap from a tracking error, so they assumed the worst and worked around the tool.
Users independently validated the data outside the platform before facing a client, because a "0" might just be a tracking error.
Account Executives had stopped sharing the auto-generated reports altogether — system errors risked reflecting on their own professional competence.
To see unbundled product footprints and contract terms, users constantly context-switched between the scorecard, the CRM system of record, spreadsheets, and personal living documents.
Back-end re-weighting could swing a score 20 points with zero customer action — leaving teams unable to explain a "win" or a "loss."
The biggest pain point is severe data distrust. I manually verify the data outside the platform before every client meeting — because a "0" could just be a tracking error.
Leadership had heard the complaints. What was missing was a body of evidence strong enough to redirect a roadmap — and specific enough for engineering to act on. I structured the program so each layer answered a different question: who are the users, where does the workflow break, what's broken in the UI, and does the fix actually work?
I ran 16 in-depth 1:1 interviews (45-minute deep dives) across the three roles that live in the tool, and turned them into working personas built on real jobs — not job titles.
Interviews surface complaints; a blueprint surfaces systems. I built an end-to-end service blueprint of the customer lifecycle across all three roles — evidence, actions, emotions, and the backstage tooling behind each step — which made the fragmentation impossible to argue with.
A screen-by-screen expert review of the dashboard produced 15 findings, each scored on a 0–4 severity scale, mapped to a usability heuristic, and paired with a concrete, low-effort recommendation and its expected impact.
Ahead of a major product release, I ran unmoderated remote usability testing with 19 participants — 8 internal operators and 11 external users new to the tool — using think-aloud protocol on interactive prototypes, plus A/B/C preference testing on two contested design decisions.
The blueprint traced the full account lifecycle — showing, at every step, which tool each role reached for, what they felt, and where the platform quietly handed the work back to them.
Routine health checks and manual tracking to cover visibility gaps — plus hunting for upsell signals in utilization metrics.
Preparing customer-facing reviews, reconciling adoption data against contracted licenses, and rewriting AI-generated decks by hand.
Modeling pricing uplifts in spreadsheets, hunting shelfware to protect the negotiation, and opening the renewal conversation 90 days out.
The blueprint became the artifact everyone pointed at — product, engineering, and leadership finally looking at the same map instead of at each other.
Each finding got a severity rating, a heuristic, a recommendation, and a stated impact — so prioritization was a conversation about evidence, not opinion.
Two design decisions were genuinely contested inside the team. Rather than let seniority settle them, we put all three variants of each in front of real users. The results were not what the team expected.
Some external users liked keeping the dashboard visible behind it. Internal teams rejected it outright — it felt disjointed from their workflow.
Captured and focused attention; the centralized call-to-action made next steps immediate and obvious.
Keeping recommendations inline preserved context but produced visual clutter and cognitive overload.
91% found the new recommendations section easy to locate (internal users averaged 15 seconds). The critical-alerts banner scored 5/5 for visibility with 100% of internal users. And both audiences endorsed the shift from generic "product adoption" to measurable business objectives, and from monthly to weekly change tracking.
The "explore feature" control scored only 66.7% discoverability — and worse, it opened an AI chat when users expected static learning material. Internal users called the repeated buttons clutter. Testing caught it before release, not after.
The most useful finding was the uncomfortable one: a subset of users disliked all three variants, calling the experience overwhelming. We reported that plainly. Option B shipped as a baseline — not as a finish line.
The scorecard's problem was never really the score. It was that the product asked people to stake their professional credibility on a number it wouldn't explain. Once we reframed "accuracy" as "explainability," the roadmap almost wrote itself — and testing the beta features taught me that the finding worth protecting is the one nobody on the team wanted to hear.
The highlighted line is a placeholder in your voice — replace it with the lesson that feels most true to you.