Your team sells nothing like the company that invented your scorecard. Every call your reps make gets graded against MEDDIC — a framework designed for committee-heavy enterprise deals — while your actual motion is product-led, channel-driven, or built on two-call velocity. Custom call scoring exists for exactly this gap: when the way you win deals doesn't match any named methodology, your scoring criteria should come from your own won-deal patterns rather than someone else's playbook.
The mismatch rarely announces itself. Instead, it shows up as quiet dysfunction — reps reciting qualification questions that buyers find bizarre, managers arguing about what "Economic Buyer identified" means for a self-serve upgrade, and scorecards that rate lost deals higher than won ones. The rubric looks rigorous. In practice, it measures fidelity to a borrowed script, not the behaviors that actually close your deals.
This guide covers both halves of the problem. First, we'll unpack what methodologies really are, how to spot rubric misfit, and how to design criteria from your own wins. Then we'll look at how modern scoring platforms let you run a fully custom rubric across every call without adding a single manual review to anyone's calendar.
A sales methodology is a set of win patterns, extracted from one company's deals and generalized into a repeatable framework. That origin matters more than most enablement teams admit. MEDDIC emerged from complex enterprise software sales, so it assumes committee buying, long evaluation cycles, and a paper process worth mapping. The framework works brilliantly when those assumptions hold.
Every other named framework carries its own embedded assumptions. BANT presumes budget-first qualification, where a prospect either has allocated money or isn't worth pursuing. SPIN assumes a discovery-heavy motion with room for extended situation and problem questioning. Challenger expects deals won by teaching the buyer something provocative about their business. Sandler builds on pain-driven mutual qualification, while SPICED centers recurring-revenue impact and GAP frames everything as the distance between a current state and a desired future state.
None of these frameworks is wrong. However, each one encodes a specific theory of why deals close — and that theory came from someone else's market, deal size, and buyer. Adopting a methodology means adopting its assumptions. When your motion violates those assumptions, the framework doesn't degrade gracefully. It keeps producing scores that look precise while measuring things that don't predict your outcomes.
Rubric misfit produces recognizable symptoms long before anyone questions the methodology itself. Watch for three patterns in particular.
First, reps start gaming criteria that don't matter. If the scorecard rewards "identified the Economic Buyer" but your buyers are individual team leads swiping a credit card, reps learn to write "EB = signer" in every note. The box gets checked, the score goes up, and nothing about the deal improves. Behavior bends toward the rubric instead of the buyer.
Second, scores stop correlating with outcomes. Deals that scored high on every framework criterion still slip or die, while scrappy deals that "failed" qualification close cleanly. When your best reps' calls score worse than your struggling reps' calls, the instrument is broken — not the people. As Harvard Business Review notes in its analysis of gen AI myths holding sales and marketing teams back, technology amplifies the process you feed it; pointing powerful scoring tools at the wrong criteria simply produces confident noise faster.
Third, coaching conversations turn into rubric litigation. Instead of discussing what the rep did on the call, the one-on-one becomes a debate about whether "Metrics" applies to a services engagement. When managers spend more energy defending the scorecard than developing the rep, the scorecard has become the problem.
Certain go-to-market motions are structurally incompatible with the big named methodologies. If you run one of these, borrowed rubrics will chafe from day one.
Research from McKinsey's growth, marketing and sales practice consistently emphasizes that commercial excellence comes from tailoring the operating model to the actual buying journey, not from importing a universal template. Consequently, teams in these motions face a choice: keep forcing a borrowed rubric, or build criteria that describe how they genuinely win.
Custom call scoring is the practice of evaluating sales conversations against criteria derived from your own winning deals, rather than against a pre-packaged methodology. Instead of asking "did the rep execute MEDDIC?", it asks "did the rep do the specific things our closed-won calls have in common?" The rubric becomes a mirror of your motion, not a template from someone else's.
In practice, custom scoring sits on top of a conversation intelligence layer that records and transcribes every call, then applies your rubric consistently across all of them. That consistency is the entire point. A custom rubric applied by hand to a random sample of calls inherits all the old problems of manual review — small samples, inconsistent graders, and week-long feedback delays. Modern call scoring software removes those constraints, which is precisely what makes fully custom criteria practical for the first time.
Importantly, custom doesn't mean casual. A good custom rubric is more disciplined than a borrowed one, because you must justify every line with evidence from your own deals. The next three sections walk through that design process.
Start with evidence, not opinions. Pull your last twenty closed-won deals and gather the call recordings behind them — first meetings, demos, and negotiation calls. This is your source corpus. Resist the urge to begin from what leadership believes good selling looks like; beliefs are exactly how borrowed rubrics got installed in the first place.
Next, extract the behaviors those winning calls share. Listen for what reps actually did: the questions they asked, the moments they quantified impact, the way they handled pricing, the specific next steps they secured. Write each observed behavior on its own line. After a dozen calls, clusters emerge — perhaps every win included the rep asking how the team handles the problem today, or naming the implementation timeline before the buyer asked.
Then pressure-test the clusters against losses. For each candidate behavior, check a handful of closed-lost calls from the same period. Behaviors present in wins and absent in losses are your signal; behaviors present in both are table stakes, not differentiators. From this filtered list, select six to ten criteria. Fewer than six leaves blind spots, while more than ten dilutes coaching focus and makes every score a blur.
The single biggest quality gap between rubrics is specificity. "Good discovery" is a vibe — two managers will score it two different ways, and a rep can't act on it. "Asked how the team handles onboarding today, quantified the hours it consumes, and confirmed the number back to the buyer" is a behavior. It either happened on the call or it didn't.
A useful pattern for writing rubric lines is asked, quantified, confirmed. Each criterion should name the observable action (asked about renewal timing), the depth marker (quantified the cost of the current workaround), and the verification step (confirmed the buyer agreed with the summary). Structuring lines this way makes scores reproducible across graders — human or AI — and makes feedback immediately actionable.
Apply the same discipline to your scale definitions. Rather than "1 = poor, 5 = excellent," anchor each level to evidence: a 2 means the rep asked but never quantified; a 4 means asked and quantified but skipped confirmation. Anchored scales turn scoring disagreements into checkable questions about the transcript, instead of debates about taste. As a result, calibration later becomes dramatically easier.
Never launch a custom rubric directly into production. Before a single live call gets scored, run the rubric retroactively across a mixed set of past calls — wins, losses, and stalled deals your team already knows the ending to. This backtest answers the only question that matters: do the scores separate the outcomes?
Look for separation first. If closed-won calls and closed-lost calls score roughly the same, your criteria are measuring noise, and you should return to the won-deal analysis before proceeding. Second, look for surprises — a lost deal that scores highly is worth investigating, because either the rubric missed something or the deal died for reasons outside rep control. Both findings sharpen the rubric.
The pilot also builds the trust you'll need at rollout. Reps who watch the rubric correctly distinguish their best calls from their worst ones stop treating scores as arbitrary judgment. We covered why consistency beats sporadic human review in our guide to AI call scoring versus manual reviews — the same logic applies doubly to a rubric you invented yourself. Evidence, not authority, is what makes a new scorecard legitimate.
A custom rubric has no external reference material — no books, no certification courses, no conference talks explaining what a 4 on "partner co-sell orchestration" means. Your managers are the only living documentation. Therefore, calibration isn't optional hygiene; it's the mechanism that keeps the rubric meaning one thing across the whole team.
Run calibration sessions before launch and on a recurring cadence afterward. Have every manager independently score the same three calls, then compare line by line. Wherever scores diverge, the disagreement points to an ambiguous rubric line — fix the wording, add an anchor example, and re-test. Over a few cycles, the rubric converges on definitions everyone applies identically.
We walked through the full session format, drift checks, and dispute-resolution process in our post on call scoring calibration, and every technique there applies directly to custom criteria. The stakes are simply higher with a custom rubric. When a borrowed framework drifts, you can appeal to the source material; when a custom one drifts, only your calibration discipline stands between the team and score inflation.
None of this is an argument against MEDDIC, or any other framework. If you sell six-figure deals into buying committees, with security reviews, legal negotiation, and a genuine paper process, MEDDIC's assumptions match your reality — and its criteria will predict your outcomes well. The framework earned its reputation in precisely that motion.
The same holds elsewhere. BANT remains a fast, honest qualification filter for teams with hard budget gates. SPIN still teaches the best discovery questioning sequence ever documented, and Challenger genuinely fits markets where buyers reward being taught. Sandler's mutual qualification protects reps in motions plagued by tire-kickers, while SPICED and GAP give recurring-revenue teams a clean impact vocabulary.
The test is never whether a framework is good. Instead, ask whether its embedded assumptions — committee versus individual buyer, long versus short cycle, discovery-led versus usage-led — describe your motion. If they do, adopt the framework and score against it wholeheartedly. If they don't, no amount of training will make a borrowed rubric predictive. Fit is the whole game.
The table below summarizes how the two approaches behave in day-to-day coaching and deal management.
| Dimension | Borrowed Rubric | Custom Rubric |
|---|---|---|
| Source of criteria | Another company's win patterns | Your last twenty closed-won deals |
| What a high score means | Rep followed the framework | Rep did what wins your deals |
| Rep behavior it drives | Checkbox recitation, criteria gaming | Repeating proven winning behaviors |
| Coaching conversation | Debates about rubric applicability | Specific gaps on specific calls |
| Score-outcome link | Weak when assumptions don't fit | Validated in backtest before launch |
| Maintenance burden | Low, but drifts from your motion | Needs calibration and periodic review |
| Best fit | Motions matching the framework's origin | PLG-assist, channel, services, velocity, vertical motions |
Neither column is universally better. The borrowed rubric wins on adoption speed and shared vocabulary, whereas the custom rubric wins whenever your motion diverges from the framework's assumptions. Choose based on fit, then commit to the maintenance the choice requires.
Everything above works on paper regardless of tooling — but applying a custom rubric to every call, consistently, is where teams historically gave up. This is the problem Rafiki AI was built to remove. Its Smart Call Scoring capability scores every recorded conversation against MEDDIC, BANT, SPIN, SPICED, GAP, Challenger, or Sandler out of the box — or against fully custom criteria you define yourself.
Defining custom criteria mirrors the design process in this guide. You write your behavior-level rubric lines, set the anchored scale, and Rafiki AI applies them to one hundred percent of calls from the moment they end. Because the same AI evaluates every conversation against the same definitions, the grader-variance problem that plagues manual custom scoring disappears. Your backtest is easy too: point the rubric at historical calls and check the win-loss separation before anything goes live.
For enablement leaders, this changes what a rubric revision costs. Updating a criterion no longer means retraining a bench of managers and waiting a quarter to see effects. Instead, you edit the rubric, rescore, and see the impact across the whole call corpus within days. Ready to test your own criteria against real calls? Start your free trial today and run your first backtest this week.
Scores that sit in a dashboard change nothing. Rafiki AI closes the loop in two directions — into your CRM and into your coaching motion.
On the CRM side, Smart CRM Sync auto-populates methodology-specific fields for every framework Smart Call Scoring supports, and it handles custom fields with the same fidelity. If your rubric tracks "partner registered before demo" or "implementation scope confirmed," those become structured CRM data on every deal, captured from the conversation itself rather than from rep memory. RevOps gets pipeline reporting built on your motion's real signals, without adding a single field to the rep's after-call admin.
On the coaching side, the Coaching Agent — one of Rafiki AI's autonomous AI agents — turns per-call scores into per-rep guidance. It aggregates each rep's scores across your custom criteria, identifies the one or two behaviors holding them back, and surfaces specific call moments that show the gap. For frontline managers, this means one-on-ones start from evidence instead of anecdote. We described how scores become a closed coaching loop in our post on AI skill scoring; with custom criteria, that loop runs on the skills your motion actually rewards.
Named methodologies are codified win patterns from someone else's sales motion. When their assumptions match yours, use them — MEDDIC for committee-heavy enterprise, BANT for budget-gated qualification, SPIN for discovery-led selling. When they don't, stop grading your team against a borrowed answer key. The symptoms are always the same: gamed checkboxes, high scores on lost deals, and coaching hours burned litigating the rubric.
The alternative is disciplined, not improvised. Pull your recent wins, extract the shared behaviors, write them as asked-quantified-confirmed rubric lines, backtest against known outcomes, and calibrate your managers before launch. That process produces custom call scoring criteria with something no borrowed framework can offer — proof, from your own deals, that the behaviors you're scoring are the behaviors that win.
Rafiki AI makes that rubric operational at full scale: Smart Call Scoring grades every call against your criteria, Smart CRM Sync writes the results into the fields your pipeline reports depend on, and the Coaching Agent converts patterns into rep-specific development plans. Your motion is unique. In 2026, your scorecard finally can be too.
Aim for six to ten criteria. Fewer than six usually means the rubric misses whole phases of your motion — discovery behaviors get covered while next-step discipline goes unmeasured, for example. More than ten creates two problems: scores blur together because no single criterion moves the total much, and coaching loses focus because every rep has a dozen "gaps" of equal apparent weight. The filter for inclusion should be evidence, not completeness. A criterion earns its slot by appearing consistently in your won-deal calls and rarely in your lost ones. If a behavior shows up everywhere regardless of outcome, it's table stakes — train it during onboarding, but don't spend a rubric line on it. Revisit the count whenever your motion changes materially, such as after adding a partner channel or a new product tier.
Yes, and hybrid setups are common in practice. Many teams run different rubrics for different segments — MEDDIC for the enterprise pod, where committee buying makes its assumptions valid, and custom criteria for the velocity or PLG-assist pod, where those assumptions break. Others keep a named framework as the deal-qualification layer in the CRM while scoring call execution against custom behavioral criteria, since qualification state and conversation quality are genuinely different measurements. The mistake to avoid is blending both into one scorecard, because mixed rubrics reintroduce the ambiguity you built custom criteria to eliminate. Platforms like Rafiki AI support this segmentation directly: Smart Call Scoring can apply MEDDIC, BANT, SPIN, SPICED, GAP, Challenger, Sandler, or custom criteria, so each team scores against the rubric that fits its motion.
Review the rubric on a quarterly cadence, and additionally after any material change to your motion — a new product line, a shift toward partners, a move upmarket, or new pricing. The quarterly review doesn't require rebuilding from scratch. Instead, rerun the core validation: check whether recent won deals still exhibit the scored behaviors and whether scores still separate wins from losses. Criteria that have stopped discriminating get rewritten or retired, and newly observed win behaviors get drafted as candidate lines. Pair each revision with a manager calibration session, because changed wording silently changes how graders score. One caution: resist revising the rubric every time a rep disputes a score. Frequent unversioned edits destroy trust and make trend lines meaningless, so batch changes into the scheduled review and communicate them explicitly.
You can start without it, but you can't scale without it. The design work — pulling wins, extracting behaviors, writing rubric lines — requires only recordings and patience, so small teams sometimes pilot a custom rubric with manual reviews. The constraint appears immediately afterward. Manual scoring covers only a sliver of calls, different managers apply the rubric differently, and feedback arrives days after the conversation, which together erode exactly the consistency a custom rubric exists to provide. Conversation intelligence platforms remove those limits by transcribing every call and applying your criteria uniformly the moment each call ends. That full-coverage consistency is what makes backtesting practical, calibration verifiable, and rep trust durable. In short: design the rubric with human judgment, then let software do the grading at scale.
Rafiki AI's conversation intelligence platform starts at $19 per seat per month with no seat minimums and no annual commitment. Start your free trial today or book a demo to see how custom call scoring built on your own win patterns transforms coaching, CRM hygiene, and forecast confidence.
Start for free — no credit card, no seat minimums, no long contracts. Just better sales intelligence.