Skip to main content
Brand Foundations
Module 6 ~20 min

Measuring Success

Learning Objectives

  • Build a Content Scorecard for a real piece using NPS + the semantic pairs that actually predict behavior change
  • Pick the right survey instrument — Pulse, Full, or Recall — for the stage you're testing
  • Run the ship-or-revise decision tree against a real scorecard, including the edge cases

Measuring Success

Compliance content is unmeasured because nobody expects it to perform. {{PROGRAM_NAME}} content is a product, and products get measured like products. This module gives you the Content Scorecard: NPS as the anchor metric, five semantic pairs that diagnose why a piece works (or doesn’t), and a three-survey decision tree that tells you ship, diagnose, or revise — in 3–4 hours per launch.

Operator note: Replace {{PROGRAM_NAME}} with your program name throughout. The NPS question wording, the five semantic pairs, the scoring thresholds, and the ship-or-revise decision tree are part of the brand measurement system and should not be modified — they’re what keeps scorecards comparable across operators. What you do customize is how you deploy the survey (embed widget, Google Forms, Typeform) to fit your stack. Narrator scripts should be recorded in the Playbook Tier 1 voice.

Program: Brand Foundations | Duration: ~20 min | Prerequisites: Customer Journey


Quick-scan index

SectionWhat it coversTime
Learning objectivesWhat you’ll know after this module
Section 1: The Content ScorecardNPS anchor metric and five semantic pairs~10 min
Section 2: Running a ScorecardThree survey instruments and the decision tree~7 min
Module test5 graded questions, 80% to pass~3 min
Key takeawaysSummary

Learning objectives

After completing this module, you will be able to:

  1. Build a Content Scorecard for a real piece using NPS + the semantic pairs that actually predict behavior change
  2. Pick the right survey instrument — Pulse, Full, or Recall — for the stage you’re testing
  3. Run the ship-or-revise decision tree against a real scorecard, including the edge cases

Why this matters: You’re the one signing things off. A piece can read beautifully in the room and still flop with players — and the only way you’ll know is the scorecard. This module gives you the numbers to make the ship call fast, and to defend it when a stakeholder who loved the draft asks why it’s going back for a revise.


Think about this: Picture the last responsible-gambling piece your team shipped. Did anyone ever ask a player what they thought of it? If the honest answer is no, hold that thought — by the end of this module you’ll have a way to find out in 3–4 hours that a campaign brief takes longer to write.


Section 1: The Content Scorecard

Reading

Think about how you already work. You A/B test email subject lines. You measure campaign NPS. You’d never ship a sportsbook promo without some way to tell whether it landed. Your {{PROGRAM_NAME}} material deserves exactly the same treatment — because Tier 1 material is a product, not a pamphlet. The reason compliance content goes unmeasured isn’t that it’s impossible to measure. It’s that nobody ever expected it to perform. We do.

Start with why this gets measured at all. {{PROGRAM_NAME}} material competes for the same player attention as your sportsbook promos, your casino loyalty emails, and your brand campaigns. It shows up in the same inbox, on the same feed, between the same two taps. So if you’d measure those with NPS, measure this with NPS too. The bar isn’t “good enough for compliance.” The bar is commercial-grade, because the thing it’s sitting next to is commercial-grade.

The anchor metric is one question. Every scorecard starts with the same line you’d put on any product survey: “How likely would you recommend this material to a friend who gambles?” Players answer on a 0-to-10 scale, and you score it the standard NPS way — the percentage of Promoters (the 9s and 10s) minus the percentage of Detractors (anyone from 0 through 6). It’s the same number you’d run for your app, your loyalty program, or your sportsbook UX, which is the whole point: one metric, comparable across everything.

What the NPS number is telling you is easiest to read as a ladder. A score of 50 or above is exceptional — players aren’t just tolerating the material, they’re sharing it on their own, which puts it in the same range as Netflix and Spotify (50–70). A score in the 30-to-49 band is strong; it matches the benchmarks for gaming and sportsbook apps (30–50), so you can ship it confidently. A score from 0 to 29 is still positive, but it’s below target — worth shipping eventually, but first run diagnostics with the Full survey to find out what’s holding it back. And anything below 0 is a problem: you have more detractors than promoters, so revise before it goes anywhere. Notice there’s no baseline you’re scrambling to beat here. Traditional compliance material is simply unmeasured — nobody ever asks players what they think of it — so the moment you put a real NPS on a piece, you already know more than the industry standard does.

NPS tells you whether a piece works. The five semantic pairs tell you why. Once you know a score is low, you need to know which part of the experience let it down — and that’s what the pairs are for. Each one is a 1-to-7 scale where players rate the material against two opposite words, and each pair watches a different dimension of quality. Forgettable–Memorable is about stickiness: will players still remember this tomorrow? Boring–Engaging is entertainment value: did it hold attention, or did it feel like homework? Confusing–Clear is comprehension: can a player understand the message on the first read? Generic–Made for me is cultural resonance: does it feel aimed at this player, or like one-size-fits-all? And the one to watch most closely is Preachy–Respectful, which measures voice register — whether the material respects the player or lectures them. That pair is the critical one, and in a moment you’ll see why it gets special powers. (If you’re wondering whether two-word scales are rigorous enough to trust, they are: the pairs come straight from validated research — semantic differential scales from Osgood and colleagues in 1957, NPS from Reichheld in 2003, and perceived message effectiveness from Dillard and colleagues in 2007. You don’t need any of that background to use them. Players score each pair, and the numbers do the talking.)

To make this concrete, picture a real piece going under the scorecard. Here’s the kind of {{PROGRAM_NAME}} material a scorecard exists to grade — and as you read the pairs above, ask the question a player would: would they rate this Memorable? Clear? Made for me?

Playbook poster “Know Your Game” — odds and the house edge explained in plain, confident language, built to commercial-marketing production quality
The kind of piece a scorecard grades: a real {{PROGRAM_NAME}} poster, not a footer disclaimer. Run it through the five pairs — would a player call it Memorable, Clear, and Made for me, or does any dimension come back soft? View in the Playbook repo →

Reading a pair score works the same way the NPS ladder did. A pair averaging 5.5 or higher is strong — that dimension is working, ship it. Anything in the 4.0-to-5.4 range is acceptable: worth improving if you easily can, but not something that blocks a launch. Once a pair drops below 4.0 you should investigate — go read the open-ended comments for that dimension and find out what players are reacting to. And below 3.0 is a red flag: the material is genuinely failing on that dimension, so it needs a revise, not a tweak.

The Brand DNA Alarm. Here’s where Preachy–Respectful earns its special status. If that one pair scores below 4.0, stop — no matter how good everything else looks. A low score there means your material may be lecturing players, and lecturing is the single biggest failure mode for literacy content. This pair can override an otherwise strong scorecard all by itself. When the alarm trips, go back to the voice principles in Voice & Tone and revise before you ship.

Narrator Script

Narrator: [transition slide] You A/B test email subject lines. You measure campaign net promoter score. Your {{PROGRAM_NAME}} material should get the same treatment. Because Tier 1 material is a product, not a pamphlet.

[pause 3s] {{PROGRAM_NAME}} material competes for the same player attention as your sportsbook promos, casino loyalty emails, and brand campaigns. If you would measure those with net promoter score, measure this with net promoter score. The quality bar is commercial-grade, not good enough for compliance.

[show nps benchmark scale] The anchor metric is one question. How likely would you recommend this material to a friend who gambles. Scored zero to ten, calculated as percent Promoters, which are nines and tens, minus percent Detractors, which are zeros through sixes. Here are the benchmarks. Fifty or above is Exceptional. Players are sharing the material organically. Thirty to forty-nine is Strong, on par with gaming and sportsbook apps. Zero to twenty-nine is positive but below target. Run diagnostics. Below zero means more detractors than promoters. Revise before shipping.

For perspective, streaming apps like Netflix and Spotify score fifty to seventy. Gaming and sportsbook apps score thirty to fifty. Traditional compliance material is unmeasured. Nobody even asks players what they think. {{PROGRAM_NAME}} targets the same range as the products it sits alongside.

[show five-pairs grid] Net promoter score tells you whether material is working. The semantic pairs tell you why, or why not. There are five pairs, each scored one to seven. Forgettable versus Memorable measures stickiness. Boring versus Engaging measures entertainment value. Preachy versus Respectful measures voice register, and this is the critical pair. Confusing versus Clear measures comprehension. Generic versus Made for me measures cultural resonance.

[pause 5s] If Preachy versus Respectful scores below four point zero, stop. Your material may be lecturing players. This is the Brand DNA alarm. It overrides an otherwise strong scorecard. Revise against the voice principles from Voice & Tone before shipping.

[pause 3s] Hold onto these two pieces — the net promoter score bands and the five-pair thresholds. In the next section they snap together into a single decision tree, and you’ll run that tree on six real scorecards yourself.


Section 2: Running a Scorecard

Reading

Now you have the two pieces — the NPS ladder and the five-pair thresholds. This section snaps them together into something you can actually run. The whole thing is three surveys feeding one decision tree, and the total effort is only 3–4 hours spread across two weeks. Here’s how each part works.

You have three survey instruments, and each one answers a different question. Reach for the one that fits the moment.

The Pulse is your default — send it after every launch. It’s deliberately tiny: three items, under a minute to complete, and you’re aiming for about 30 responses. It asks the NPS question (0–10), the single Forgettable–Memorable pair (1–7), and one open-ended prompt for the one thing the player would change. That’s enough to give you a first read on any piece without asking much of anyone.

The Full survey is your diagnostic — reach for it when the Pulse flags something. It’s longer: seven items, around two minutes, again targeting 30 responses (or 60 if you’re A/B testing two versions of the same piece). On top of the NPS question it carries all five semantic pairs (1–7) plus open-ended feedback, so when a score comes back soft, the Full survey is what pinpoints exactly which dimension is dragging it down.

The Recall survey is your retention check — it goes out 7 to 14 days after players saw the material. It’s four items and works fine with about 25 responses, because here you’re reading open-ended patterns rather than precise averages. And it tests three things neither of the other surveys can touch: free recall (“What do you remember?”), true/false knowledge checks, and behavioral follow-through — literally asking whether the player did anything differently. That last one is the real prize. A piece can score beautifully on the Pulse and still fail on Recall, and when it does, you’ve learned something important: the material was engaging in the moment but didn’t actually stick.

A quick word on those response targets, because they’re not arbitrary. Aim for 30 responses on the Pulse and Full for the numbers to be reliable. If you’re A/B testing two versions, you need 60 — 30 per variant — so each side stands on its own. The Recall can run lighter at 25 precisely because you’re looking for patterns in what people say, not a tight average.

Now the decision tree — this is the part you run after every launch, and it’s only three branches. You always start the same way: send the Pulse. Once the results land, the score routes you to one of three calls. If NPS is 30 or above and every pair clears 4.0, you ship — the material is performing, so launch it and schedule a Recall in 7–14 days to confirm it sticks. If NPS lands in the 0–29 band or any single pair dips below 4.0, you diagnose — something’s off, so run the Full survey and let its five pairs tell you which dimension needs the work. And if NPS is below 0 or Preachy–Respectful is below 3.0, you revise — stop shipping, take it back to the Voice & Tone principles, fix it, and retest. Notice that the ship branch needs both conditions true, while the diagnose and revise branches trip on either one — that’s what keeps a single weak dimension from sneaking out the door behind a healthy NPS.

It’s worth walking two of these to see the tree think. Say your team launches a myth-busting campaign about slots RNG — something like the card below — and the Pulse comes back from 35 responses with an NPS of +38 and Memorable at 5.8.

Playbook myth card “The hot streak myth” — explains in plain language that past spins don't change the odds of the next one, with a large orange 0% stat for the chance past results affect the next spin, on a navy background with the PlayBOOK wordmark
The piece behind the numbers: a real myth-busting card on the “hot streak” distortion. This is what an NPS of +38 and a Memorable of 5.8 are actually scoring — and what a Recall survey would test a week later to see if the math stuck. View in the Playbook repo →

Both clear their bars comfortably — 38 is in the strong band, 5.8 is above 5.5 — so the call is obvious: ship it, and schedule the Recall in about ten days. Now run the same piece with a different result: NPS +14 and Memorable 4.2. The 14 is positive but well under the 30 you need, so you don’t ship and you don’t revise blind — you diagnose. You send the Full, and it comes back with Preachy–Respectful at 3.6 and Generic–Made for me at 3.4. There’s your answer: the material is lecturing and it doesn’t feel targeted. The Brand DNA Alarm has already tripped on that 3.6, so the path forward is clear — revise the voice, add segment-specific framing, and retest before this goes anywhere.

Across the life of a piece, the three instruments settle into a rhythm. The Pulse fires on every launch — that one’s non-negotiable, it’s your first read. The Full comes out whenever the Pulse flags a problem and you need to know which dimension to fix. The Recall runs 7–14 days after launch for your flagship material, to confirm it stuck. And for your highest-traffic pieces, run a Full again quarterly as a drift check — material that scored well a year ago can quietly slip as the audience and the context move on. Add it all up and a single launch costs about 3–4 hours over two weeks: roughly an hour to set up and send the Pulse, an hour to collect and read it, an hour to set up the Recall, and an hour for the final read. Less time, in other words, than it takes to write one campaign brief.

One last practical thing: how you actually put the survey in front of players. {{PROGRAM_NAME}} ships a ready-to-use survey widget that auto-calculates NPS and every semantic-pair score and color-codes the results for you, and you’ve got a few ways to deploy it depending on your stack. The recommended route is the embed widget — about five minutes to drop an iframe onto any page, with the auto-scoring and export built in. If you’d rather stay in familiar territory, Google Forms takes around ten minutes and is free, though you’ll score it by hand in a spreadsheet. And if you want a more polished experience with analytics already wired up, Typeform or SurveyMonkey runs about fifteen minutes to set up, though you may need a paid plan. The method matters less than the discipline around it, which is the point of the integration note below.

Operator integration point: Pick one deployment method and make it the team default before you run your first scorecard — comparisons only hold when every piece is scored the same way. Whichever you choose, route results to one shared place (a tracker or dashboard) so quarterly drift checks have history to read. If your organization has its own survey platform, privacy rules, or data-retention policy for player feedback, link it here: Internal Policies

Narrator Script

Narrator: [transition slide] Three surveys. One decision tree. The entire process takes three to four hours spread across two weeks.

[show three-instruments card] The Pulse is your default. Three items, under a minute, target thirty responses. It gives you the net promoter score, the Forgettable versus Memorable pair, and one open-ended question. Send it after every launch.

The Full is your diagnostic. Seven items, about two minutes, target thirty responses or sixty if you are A/B testing two versions. It adds all five semantic pairs so you can pinpoint which dimension needs work. Run it when the Pulse flags a problem.

The Recall goes out seven to fourteen days after players saw the material. It tests three things the other surveys cannot measure. Free recall, which is simply asking what do you remember. True-or-false knowledge checks. And behavioral follow-through, which asks did you actually do anything differently. Material that scores well on Pulse but fails on Recall is engaging but not memorable.

[show decision tree] The decision tree is straightforward. Step one, send the Pulse. Step two-A, if net promoter score is thirty or above and all pairs are four point zero or above, ship it. Schedule the Recall survey in seven to fourteen days. Step two-B, if net promoter score is zero to twenty-nine or any pair is below four, run the Full to diagnose. Step two-C, if net promoter score is below zero or if Preachy versus Respectful is below three point zero, stop shipping and revise.

[pause 3s] Sample sizes matter. Target thirty responses for reliable results. If you are comparing two versions of the same material, you need sixty, which is thirty per variant.

[show deploy-options card] {{PROGRAM_NAME}} ships a ready-to-use survey widget that auto-calculates scores with color-coded results. You can embed it as an iframe in about five minutes. Or use Google Forms in about ten minutes with manual scoring. Or use Typeform or SurveyMonkey in about fifteen minutes for a polished experience. Whichever you pick, make it the team default so every scorecard stays comparable.

[transition to exercise] Now put the tree to work. Six real scorecards are waiting below — for each one, make the call: ship, diagnose, or revise. The thresholds you just learned decide every one of them, no gut feel allowed. Work through them, then continue.

Exercise: Ship, Diagnose, or Revise?

Type: scenario

Instructions: Six scorecards just landed on your desk. For each result, make the call the decision tree would make: Ship, Diagnose (run the Full survey), or Revise. Use the thresholds from this lesson — no gut feel allowed.

ItemCorrect CategoryFeedback CorrectFeedback Incorrect
Pulse survey, 32 responses: NPS +35, Forgettable-Memorable 5.7.ShipCorrect — this one ships. NPS 35 clears the 30+ threshold and Memorable 5.7 clears 4.0 with room to spare. Schedule the Recall survey in 7-14 days to confirm it sticks.Run the numbers again. NPS 35 is above the 30+ shipping threshold and Memorable 5.7 is above 4.0 — nothing here triggers diagnostics or revision. Ship it and schedule the Recall survey in 7-14 days.
Pulse survey, 31 responses: NPS +21, Forgettable-Memorable 4.6.DiagnoseCorrect — positive but below target. NPS 21 sits in the 0-29 band, so the decision tree says run the Full survey. All five semantic pairs will pinpoint which dimension is dragging the score.NPS 21 is positive, but the shipping bar is 30+. The 0-29 band means run the Full survey to diagnose — it isn’t a crisis, so revising blind would be premature.
Pulse survey, 30 responses: NPS -8, Forgettable-Memorable 5.1.ReviseCorrect — stop shipping. NPS below 0 means more detractors than promoters, and a decent Memorable score doesn’t buy that back. Revise against the Voice & Tone principles and retest.A negative NPS is the hard stop. Below 0 means more detractors than promoters — the decision tree says revise and retest, not diagnose. A decent Memorable score doesn’t override it.
The Pulse looked fine, so you ran the Full survey, 33 responses: NPS +44, but Preachy-Respectful comes back at 2.8.ReviseCorrect — the Brand DNA alarm just went off. Preachy-Respectful below 3.0 overrides an otherwise strong scorecard, NPS 44 included. The material is lecturing players — revise the voice before it ships.NPS 44 looks great, but the Brand DNA alarm overrides it. Preachy-Respectful at 2.8 is below the 3.0 red-flag line, which means the material is lecturing players. Stop shipping and revise the voice.
Pulse survey, 34 responses: NPS +33, Forgettable-Memorable 3.8.DiagnoseCorrect — the NPS clears the bar, but Memorable doesn’t. Any pair below 4.0 sends you to the Full survey, even with NPS above 30. Diagnose before you ship a piece nobody remembers tomorrow.A strong NPS doesn’t ship on its own. The rule is NPS 30+ and all pairs at 4.0+ — Memorable at 3.8 fails the second half, so run the Full survey to find out why it isn’t sticking.
Pulse survey, 40 responses: NPS +52, Forgettable-Memorable 6.1.ShipCorrect — that’s exceptional territory. NPS 50+ means players are sharing the material organically, on par with streaming apps, and Memorable 6.1 clears every threshold on the card. Ship it and schedule the Recall.There’s nothing to fix here. NPS 52 is in the Exceptional band and Memorable 6.1 clears every threshold on the card — ship it. The only follow-up is a Recall survey in 7-14 days.

Module Test

Narrator: [transition slide] Time to test what you have learned. Five questions on the Scorecard framework — the net promoter score bands, the five semantic pairs, the three instruments, and the decision tree. You need eighty percent to mark this module complete. Take your time. If you do not pass on the first try, review the material and retake the test.

Question 1

Assesses: Learning objective 1

Stem: What does the NPS question in the Content Scorecard ask players?

OptionText
A”Did you find this material useful?”
B”How likely would you recommend this material to a friend who gambles?”
C”Would you read this material again?”
D”How satisfied are you with our responsible gambling program?”

Correct: B

Explanation: The NPS question is: “How likely would you recommend this material to a friend who gambles?” It’s the same recommendation metric used for commercial products — because Tier 1 material is a product.

Source: Research Methodology


Question 2

Assesses: Learning objective 3

Stem: What is the minimum NPS score needed to ship without further investigation?

OptionText
A10
B20
C30
D50

Correct: C

Explanation: NPS 30+ with all semantic pairs at 4.0+ means you can ship. This matches the “Strong” NPS range for gaming and sportsbook apps — because your literacy material competes for the same player attention.

Source: Research Methodology


Question 3

Assesses: Learning objective 1

Stem: Which semantic pair triggers the “Brand DNA alarm” when it scores below 4.0?

OptionText
AForgettable — Memorable
BBoring — Engaging
CPreachy — Respectful
DGeneric — Made for me

Correct: C

Explanation: Preachy-Respectful below 4.0 is the Brand DNA alarm. It means your material is lecturing players — the single biggest failure mode for literacy material. Revise against voice principles (Voice & Tone) before shipping.

Source: Brand Personality


Question 4

Assesses: Learning objective 3

Stem: Your Pulse survey returns NPS 22 and Memorable 5.2. What’s the next step?

OptionText
AShip it — 22 is a positive NPS
BRun the Full survey to diagnose which dimensions need work
CSend the Recall survey to check retention
DRevise the material immediately — this is a failing score

Correct: B

Explanation: NPS 22 is below the 30+ threshold for shipping. It’s not a crisis (NPS is positive), but it needs diagnostics. Run the Full survey with all five semantic pairs to identify exactly which dimensions are dragging the score down.

Source: Research Methodology


Question 5

Assesses: Learning objective 2

Stem: What does the Recall survey measure that Pulse and Full surveys do not?

OptionText
AWhether players enjoyed the material in the moment
BWhich semantic pairs scored lowest
CWhether players remember key messages and changed behavior 7-14 days later
DHow many players completed the survey

Correct: C

Explanation: The Recall survey goes out 7-14 days after exposure. It tests three things the other surveys can’t: free recall (what do you remember?), knowledge retention (true/false checks), and behavioral follow-through (did you actually do anything differently?). Material that scores well on Pulse but fails on Recall is engaging but not memorable.

Source: Research Methodology


Key Takeaways

  • Tier 1 material is a product, not a pamphlet — measure it with the same rigor as your entertainment content
  • NPS is the anchor metric: “How likely would you recommend this material to a friend who gambles?” — target 30+ to ship
  • Five semantic pairs diagnose why material works or doesn’t: Forgettable-Memorable, Boring-Engaging, Preachy-Respectful, Confusing-Clear, Generic-Made for me
  • Brand DNA Alarm: Preachy-Respectful below 4.0 is an automatic hold — regardless of other scores
  • Three instruments: Pulse (every launch), Full (diagnostics), Recall (7-14 day retention check)
  • Decision tree: NPS 30+ and all pairs 4.0+ = ship; NPS 0-29 or any pair below 4.0 = run Full; NPS below 0 or Preachy-Respectful below 3.0 = revise
  • Total effort: 3-4 hours across two weeks per launch — less time than a single campaign brief

Source basis

You don’t need to read these but they’re published research related to the scorecard — in case a stakeholder asks.

  • Osgood, C.E., Suci, G.J., & Tannenbaum, P.H. (1957). The Measurement of Meaning. University of Illinois Press. Why it’s here: the original semantic differential methodology — the academic basis for using bipolar 1-7 pairs (Preachy-Respectful, Boring-Engaging, etc.) to measure perceived qualities of a stimulus.
  • Reichheld, F.F. (2003). The one number you need to grow. Harvard Business Review, 81(12), 46-54. Why it’s here: the foundational NPS paper — the methodology and the recommendation-to-a-friend question wording the scorecard borrows directly.
  • Dillard, J.P., Shen, L., & Vail, R.G. (2007). Does perceived message effectiveness cause persuasion or vice versa? 17 consistent answers. Human Communication Research, 33(4), 467-488. Why it’s here: evidence that perceived message effectiveness (what the semantic pairs measure) actually predicts downstream behavior change — not just reported feelings.

References

ResourceWhat to use it for
Research MethodologyFull NPS and semantic pair methodology, survey templates
Brand PersonalityVoice principles for revising Preachy-Respectful failures
Survey WidgetEmbeddable scorecard survey with auto-scoring

Was this module helpful?