Case Study — HIPAA-Compliant Healthcare Forms SaaS

40,000+ support tickets. 9-stage NLP pipeline. Support noise turned into a product roadmap.

How a healthcare forms platform turned unstructured Zendesk support tickets into classified, prioritised product intelligence — with PHI redacted first, on AWS.

Stack AWS SageMaker Zendesk Python AWS Comprehend
40,000+
Support tickets analysed end-to-end
9
Pipeline stages from raw fetch to segmentation
25
Jobs-to-Be-Done categories classified against
750+
Lines of production classification code delivered

Before.

The support team had tens of thousands of Zendesk tickets and no way to turn them into product decisions. Each ticket was free text: mixed sentiment, messy feature references, PHI mixed in.

"Users report problems" was the extent of the signal. No prioritisation, no grouping by product area, no link to the jobs users were actually trying to do.

The Situation
  • Tens of thousands of support tickets as unstructured free text
  • PHI (patient data) embedded in ticket bodies — a compliance risk
  • No grouping by feature or product area
  • No way to tell which issues mattered most
  • Product roadmap decisions made without ticket evidence

What we did.

Built a 9-stage NLP pipeline on AWS that redacts PHI first, then classifies every ticket into product intelligence.

Step 1-2 — Zendesk Fetch + S3 Caching
Pulled tickets from the Zendesk API, cached raw JSON to an encrypted S3 bucket (KMS), verified round-trip reads. Secrets in AWS Secrets Manager.
Step 3 — PHI Redaction (compliant, multi-layer)
Redacted protected health information from ticket bodies before any analysis — AWS Comprehend PII detection, Presidio, keyword layers, and a redaction tracker. Healthcare means redaction is step one, not an afterthought.
Step 4 — Root-Cause Extraction
Extracted which FormDR feature or product each ticket is really about (signature field, form submission, EHR sync, etc.) with sentiment attached. "The signature part of the form fails" became Signature / Collect / negative.
Step 5 — JTBD Classification
A 750-line classifier mapped tickets across the full base to 25 Jobs-to-Be-Done (12 functional, 8 emotional, 5 social), with confidence scores. "Eliminate Manual Data Entry" ranked top with an 90/100 opportunity score.
Step 6-7 — KANO + Topic Modelling
Classified features by KANO category (must-be, performance, delight) and ran topic modelling to surface the underlying issue clusters across tickets.
Step 8-9 — Visualisations + Segmentation
Generated decision-ready charts and segmented tickets by customer and issue type, so the team could see who was affected and how often.

Results.

40,000+
Tickets classified end-to-end from noise to product intelligence
25
JTBD categories — support volume mapped to user jobs
90/100
Top opportunity score: "Eliminate Manual Data Entry" (EHR sync)
100%
PHI redacted before any analysis touched the data

From "18 users report EHR problems" to a defensible recommendation: prioritise EHR API reliability and add integration health monitoring — with the ticket evidence behind it.

What this proves.

NLP & Text Classification

Sentiment, root-cause extraction, JTBD classification, KANO, and topic modelling on real support data — built, not demoed.

HIPAA / PHI Handling

Multi-layer redaction on AWS before analysis — compliance is part of the pipeline, not a bolt-on.

From Text to Roadmap

Unstructured support volume became a prioritised product roadmap with per-ticket evidence.

Jake McMahon
Jake McMahon
ProductQuant

10 years building growth systems for B2B SaaS companies at $1M–$50M ARR. BSc Behavioural Psychology, MSc Data Science. This engagement required architecting a multi-branch retention system in Chameleon that synthesized PostHog behavioral data with real-time billing triggers to intercept churn at the point of intent.

What this looks like for your company

Chameleon Churn Prevention Flows.

A retention interception system built in Chameleon — triggered by PostHog behavioural signals, branching by cancellation reason, and measured from day one.

  • 3–5 churn signals identified and validated with trigger thresholds defined
  • 3 intervention flows designed and live: low-engagement nudge, feature discovery, CS escalation
  • Message copy, audience targeting, display conditions, and frequency caps configured
  • PostHog integration: flows triggered by behavioural segments
  • Post-Chameleon win-back sequence design for accounts the interception flow doesn’t save
$1,997 · 10 days
Right for you if
  • Using Chameleon (or evaluating it) for in-product messaging
  • Cancellation rate visible in your data but no systematic interception at the point of intent
  • CS team finding out about churn after the fact rather than before the decision is made

You find out when they cancel.

A 15-minute call is enough to know whether what we do is relevant to where you are. No pitch. Just a conversation about your specific situation.