Research case study · T-Mobile
NORA. Turning root-cause analysis into answers a care agent can use.
NORA delivers findings from T-Mobile’s root-cause analysis (RCA) engine to care agents. It checks a customer’s account, device, and network and reports what’s likely wrong. The findings were accurate — but written for higher-tier care agents, not a frontline agent with a customer on the line.
We’re adding a persona-aware layer that turns NORA’s technical output into fast, easy-to-understand diagnostics and troubleshooting steps.
- Role
- Senior Experience Designer
- Team
- Project manager (my side) · 2 software engineers (T-Mobile)
- Timeline
- May 2026 – present · Phase 2 completed September 2026
- Methods
- Generative interviews · Contextual inquiry · Personas · Journey mapping · Concept validation · AI output evaluation
- Tools
- Dovetail · Figma · Figma Make · Claude · OpenAI API
- Where it stands
- The consumer care agent prompt is built, tuned against the Playbook, and ready for testing in a controlled NORA environment.
The outcome in one line
Before
“The exact root cause of the customer’s issue is that their tethering data usage has exceeded the 5GB allocation available within their subscribed plan (SOC: [code] among others)…”
After
“The customer’s tethering data has reached the plan limit, triggering speed restrictions.”
Same diagnosis, rewritten so an agent can read it and say it mid-call. See how the Playbook rules got it there →
A two-part research effort
Part 1 · User research
Build the foundation
Interviews, observation, synthesis, and a validated persona became the principles and rules of the NORA Playbook.
Protocol · 10 interviews · 5 observations · Persona · Playbook
Part 2 · AI research
Validate the prompt
The Playbook became a system prompt, then outputs were validated with SMEs and agents and tested at scale against the rules.
System prompt · Concept validation · Evaluation harness
Part 1
User research
Understand the agents, their calls, and what a first response has to do — then write the rules.
01 · The problem
Accurate answers, hard to use.
T-Mobile’s engineers came to us with a clear signal: care agents found NORA’s output too complicated. We wanted to understand why before changing anything.
“The summarization and the next steps and action in the output, it’s still way too technical.”
- Five sectionsof technical findings to read while the customer waits.
- Signal metrics and site IDsRSRP in dBm, RSRQ, a cell site code.
- Next steps written for engineers“Check localized coverage… in site monitoring tools.”
- No customer-ready wordingThe agent still has to translate all of it, live.
02 · Research plan
Start with the decision, not the answer.
The engineers’ feedback told us that agents struggled. The research had to tell us why — and what a better first response would need to contain. One constraint shaped everything: no new UI, no builds. Every improvement had to come from what NORA says, not from changes to the screen.
- 01
Desk research
Frontline care work, agent cognitive load, and how AI assistance is used in care.
- 02
Stakeholder discovery
Care SMEs, engineers, and consumer care leads on how NORA is used today.
- 03
Agent interviews
10 remote 1:1 sessions with frontline agents.
- 04
Observation
5 sessions watching agents use NORA as they worked.
03 · Interview protocol
Writing questions that don’t lead.
Each session followed the same path — starting from real work, moving toward NORA.
- 1
Walk me through a recent call
What was difficult? What was easy?
- 2
Expectations
When you submit a query to NORA, what do you expect to see?
- 3
Translation
How do you explain the issue to the customer on the phone?
- 4
Escalation
How do you handle calls where a ticket needs to be created?
- 5
NORA output
What’s difficult to interpret? What’s easy?
The bias
SMEs had already told us agents found the output hard to read. My first draft took that as fact — some questions were written to confirm it.
AI-assisted vetting
I ran the draft through Claude to check for leading and compound questions, bias, and gaps. It flagged the assumption.
The fix
Balanced questions — what’s difficult and what’s easy — so agents could tell us if NORA worked. T-Mobile SMEs then reviewed the protocol for relevance.
04 · Interviews & observation
Listening first, then watching.
10
remote 1:1 interviews with T0/T1 agents
45–60
minutes per interview
5
one-hour observation sessions
1
tool to record, transcribe, and tag: Dovetail
What observation showed that interviews didn’t
What we saw
Agents asked NORA only once — no follow-up questions.
What it meant
The first response has to do all the work: diagnosis, next steps, and talk track in one pass.
What we saw
After entering caller details, agents typed very basic prompts.
What it meant
NORA can’t rely on a well-written prompt. The output has to be useful no matter how the agent asks.
What we saw
Ticket creation made agents anxious.
What it meant
Escalation needs a clear, confident path: when to file, what to say, what happens next.
05 · Synthesis
From transcripts to principles.
I tagged moments across all 10 transcripts in Dovetail to find recurring themes. AI sped up the clustering, with one rule: for every theme it proposed, I asked for the supporting quotes and timestamps and checked them against the transcripts. No evidence, no theme.
Translation is the work
Agents have to interpret technical signals before they can explain them with confidence.
Live calls demand scannability
A short headline, a clear next step, and speakable wording beat comprehensive detail.
Confidence needs visible limits
Agents need to tell a reliable diagnosis from an informed guess.
Consistency protects the experience
The same root cause should produce the same explanation and action, whichever agent takes the call.
Empathy must survive the handoff
Guidance should acknowledge the customer’s situation, not just report system status.
Seven guiding principles
I wrote seven principles that became the foundation of the NORA Playbook.
Research signal
Principle
Translation is the work
01Speak Plainly
Live calls demand scannability
02Surface Only What Moves Things Forward
Agents asked NORA only once (observation)
03Always Land on a Next Action
Ticket anxiety (observation) + empathy
04Make the Handoff Clean
Consistency protects the experience
05Use Consistent Structure
Live calls demand scannability
06Balance Skimming and Reading
Confidence needs visible limits
07Communicate System Status
Design implication
Every first response pairs a short diagnosis with confidence cues, customer-ready language, and one next best action the agent is allowed to take.
06 · Persona
One persona: the Frontline Consumer Care Agent.
Interviews and observation showed T0 and T1 agents share nearly all the same goals and frustrations. T1 agents take messier, higher-stakes calls, but the core need is the same — so we built one persona, validated by T-Mobile’s SMEs.
Who they are
Early-career, often in their first 6–12 months. Comfortable with consumer tech, not RF engineers. Ping peers before escalating.
What they want
Resolve on first contact. Keep calls steady. Give clear, confident, consistent answers — without guessing.
What gets in the way
Output that reads like engineering signal. 2–5 possible actions instead of one, so agents freeze or escalate.
The principle we designed by: treat the care agent like a customer.
07 · The Playbook
Turning research into rules an AI can follow.
The Playbook is the single source of truth for everything NORA says to a consumer care agent. I built it as an interactive reference in Figma Make, on top of T-Mobile’s existing Core Communication Standards — Clear and Accurate · Empathetic and Protective · Respectful of Time. Additions and edits went to T-Mobile care SMEs in bi-weekly review meetings, so the rules were checked against real care work as they took shape.
Every NORA output should feel like a briefing from a knowledgeable colleague who knows what you need and respects your time — not a system log.
Capabilities & guardrails
NORA informs the agent. It doesn’t decide, act, or replace judgment — and it says when it has gaps.
Target persona
The Frontline Consumer Care Agent, as the lens for every rule.
7 core principles
The foundation every section rule builds on.
Audience styles
Headline style for the agent (“Reset network settings”); script style for the customer (“Go ahead and restart your gateway…”).
Controlled vocabulary
Preferred and avoided terms, kept as a living list.
Section rules
Structure, length, tone, and hedging rules for each part of the output, with do/don’t examples.
Tone & personalization
Five tones matched to the customer’s situation.
One structure for every response
1 · Findings
One-line summary, then Account, Device, and Network diagnostics — only what’s relevant.
2 · Next steps
Ranked actions the agent can take, with optional wording for the customer.
3 · Last resort
If nothing works: exactly what ticket to file and what to tell the customer.
One finding, before and after
Before
“The exact root cause of the customer’s issue is that their tethering data usage has exceeded the 5GB allocation available within their subscribed plan (SOC: [code] among others). In alignment with the provisioning policy and SOC restrictions attached to this account, throttling is applied after the tethering threshold is breached…”
Three sentences of policy before the diagnosis. Internal codes.
After
“The customer’s tethering data has reached the plan limit, triggering speed restrictions.”
Leads with the finding. Plain language. One sentence.
From Playbook to prompt. NORA runs on GPT models, so the rules had to become instructions a model could follow. I translated the Playbook into a structured markdown system prompt — principles, section rules, suppression lists, and tone logic.
08 · The pushback
Should agents label a caller’s sentiment by hand?
The proposal
Agents identify the caller’s sentiment and add it to their prompt, so NORA responds in the right tone.
- Bias. Two agents would label the same caller differently.
- Discoverability. Nothing in the interface would tell agents the option existed — and under a no-build constraint, nothing could be added.
- Research quality. Free-form sentiments can’t be compared, tested, or improved.
My solution
Signal-based tone triggers — no manual flagging.
NORA reads signals it already has, like repeat callers, new customers, and time of day…
…maps them to five predefined tones…
…and writes tone-appropriate scripts automatically.
No extra work for agents, no labeling bias, and fixed categories we can test. It became the Playbook’s tone and personalization framework.
Part 2
AI research
Turn the Playbook into a prompt, then find out whether the model actually follows it.
09 · Validating the new outputs
Showing people the new NORA before anything ships.
With nothing to build, I used the system prompt and screenshots of real NORA outputs to generate new example responses in Claude, then shared them with T-Mobile care SMEs and frontline agents — in meetings, by email, and in shared documents.
The reaction was overwhelmingly positive. One issue surfaced: agents were confused when a category came back “Incomplete” even though one of its checks had failed.
The fix: a status precedence rule
If any check fails, the category is Bad — even if others didn’t complete. If nothing failed but something couldn’t be checked, it’s Incomplete. Only a category where everything completed and passed is Good. An agent never sees “Incomplete” hiding a real problem. The rule went into the Playbook, the prompt, and the evaluation harness.
10 · Evaluation
A rule the model follows three times out of five isn’t a rule yet.
You can’t see that from a single output. So I built an evaluation harness with Claude Code that runs the system prompt through the OpenAI API at scale.
- 1
Scenarios
Real NORA diagnostic outputs, plus synthetic cases for situations the real ones don’t reach.
- 2
Repeat runs
Each scenario runs several times through the model.
- 3
Automated checks
39 machine-verifiable rules: section structure and order, one status per category, summary length, approved scripts only, and leak detection for every account number, device ID, site code, and internal system name in the source data.
- 4
Dashboard
A plain-language page showing which rules broke, where, and whether they broke every time or only sometimes.
- 5
Tune & re-run
Adjust the prompt where rules break unpredictably, then run again.
What it can’t check
Judgment calls — whether a diagnosis is right, or whether a suggested step makes sense. A clean run means no automated rule broke, not that the output is correct. Those questions still need SME review. Tone scenarios are being added so all five tones are validated.
11 · Learnings & next
Standards only matter if you can measure them.
- 1
Design for the moment of use
A rep mid-call, with the customer waiting, set the length, order, and tone of every output.
- 2
Check your own assumptions first
Stakeholders told me the output was hard to read. My protocol had to leave room for agents to tell me otherwise.
- 3
Observation reveals what interviews miss
“Agents ask only once” never came up in an interview — and it shaped the entire output structure.
- 4
With no build budget, language is the design
Every improvement came from what NORA says, not what the screen shows.
- 5
Test AI like a system, not a demo
Consistency across repeated runs shows which rules the model actually follows.
What’s next
Keep refining
Continue tuning the consumer prompt after testing in a controlled NORA environment.
Scale the approach
Apply what we learned — and the materials we built — to a Playbook and prompt for three user groups in T-Mobile for Business.