Omnichannel HCP Personalization & Next-Best-Action Platform — Architecture
Healthcare commercialisation decision intelligence

From approval to adoption,
every HCP interaction has to count.

Commercial teams must convert clinical evidence, approved content, market-access context and omnichannel engagement signals into one timely, compliant and explainable decision.

Omnichannel HCP Personalization
& Next-Best-Action Platform

An end-to-end machine-learning system that determines which HCP to engage, what action to recommend, which approved content to use, through which channel and at what time — and measures whether the action created incremental value.

End-to-end machine-learning system · architecture specification
The objective is not more commercial activity.
It is better allocation of commercial effort.
The approval-to-adoption gap Clinical evidence and regulatory approval are followed by a gap covering market access and HCP engagement, bridged by the next-best-action platform, leading to appropriate patient adoption. Clinical evidence trials, publications, real-world data Regulatory approval label, indication, MLR-approved content THE GAP · WHERE VALUE IS WON OR LOST Market access formulary, reimbursement HCP education and engagement SIX COMMERCIAL INPUTS HCP profile and affiliation Approved content and MLR status Access barriers Consent and suppression Channel history Recent engagement and intent Next-Best-Action Platform retrieve · rank · adjust for utility · enforce eligibility explain · measure incrementality ONE RANKED, ELIGIBLE AND EXPLAINABLE ACTION Appropriate patient adoption the outcome the commercial system exists to serve Approval does not produce adoption. The gap is crossed one HCP decision at a time.

Where commercial value is lost

Speed to adoption

Once a therapy is approved, commercial teams have a limited window to establish awareness, communicate the evidence, address access barriers and build appropriate adoption. Slow or poorly coordinated execution weakens launch momentum.

Lost launch velocity

Relevance at scale

HCP profiles, content, consent, access, campaign and engagement data sit across disconnected systems. The right HCP can therefore receive the wrong message, through the wrong channel, at the wrong point in the journey.

Fragmented decision-making

Commercial efficiency

Field capacity, marketing budgets and HCP attention are finite. Every unnecessary visit, repeated email or underused content asset consumes budget and increases fatigue and opt-out risk.

Wasted commercial effort

One connected decision for every planning cycle

HCP×Action×Approved Content×Channel×Time Window

Inputs

  • HCP profile
  • Commercial context
  • Approved content
  • Access barriers
  • Consent
  • Channel history
  • Recent intent
  • Field capacity

Decision process

  1. Retrieve
  2. Rank
  3. Adjust for business utility
  4. Enforce eligibility
  5. Explain
  6. Measure incrementality

Optimisation objective

Expected incremental value
− fatigue
− opt-out risk
− operational cost

Mandatory constraint

Approval, effective date, segment, channel permission, consent and contact-frequency compliance.

This is not another campaign dashboard. It is a decision system for allocating constrained commercial resources to the interactions most likely to create incremental value.

Domain terminology used in the architecture
HCP
Healthcare professional — the prescriber this system recommends to.
HCO
Healthcare organisation — the hospital, clinic or practice an HCP is affiliated to.
KeyMessage
The smallest approved unit of content: a single claim, a safety block, one data figure.
ContentModule
A group of key messages approved to appear together.
Campaign
The campaign-level container that delivers content modules through a channel.
MLR
Medical, legal and regulatory review. Approves content and sets its expiry date. Authoritative — the platform reads its decisions and cannot override them.
NBA
Next-best-action — the decision this platform produces.
Eligibility
The deterministic decision that an HCP–action–content–channel–window combination satisfies approval, effective date, segment, channel permission, consent and contact-frequency rules, and is therefore servable. Distinct from relevance, and evaluated separately.
Channels FIELD_DETAILBROADCAST_EMAILREP_EMAILTHIRD_PARTY_MEDIAWEB_PORTALVIRTUAL_MEETINGEVENT
Segments EarlyAdopterGrowthPotentialEstablished
Approval states ApprovedInReviewExpired — only Approved is servable
From commercial problem to operating boundary

Before a model can recommend an action, the platform must know which systems provide the evidence, who owns the objective and which decisions it is never permitted to override.

View 01L0 · Context
Commercial operating boundary

One governed decision loop from approved content to HCP engagement

The platform sits between content governance, commercial strategy, field execution and HCP response. It does not decide what is allowed, and it does not replace the field team — it turns fragmented signals into one ranked, eligible and explainable action.

What matters is not more outreach — it is coordinated, compliant and measurable commercial action.

Approved evidence Commercial objective Decision intelligence ranked · eligible · explainable Field execution HCP engagement The platform's job is to connect approved evidence and business objectives to field execution and measured HCP response.
Inputs and authorities Decision intelligence Activation, response and governance Brand & Content Operations Campaigns · content modules approved key messages Supplies approved content inventory MLR & Content Governance Approval state · effective date expiry · channel permission AUTHORITATIVE Supplies authoritative eligibility state Field Representative & CRM Territory context · field capacity planning-cycle request Requests plan · executes in the field Approved content and metadata Authoritative approval and eligibility state Planning request and field context Ranked eligible actions and rationale DECISION INTELLIGENCE Omnichannel HCP Personalization & Next-Best-Action Platform One ranked, eligible and explainable recommendation per planning cycle 1 · Retrieveeligible action space 2 · Rankengagement value 3 · Adjust for utilityfatigue · cost 4 · Verify eligibilitynon-overridable 5 · Explainsignals · proof HCP × Action × Approved Content × Channel × Time Window The model prioritises only within the permitted action space. It never authors content or overrides approval. Approved action through permitted execution systems Engagement and outcome signals delayed · hours to days Performance · uplift and suppression Objectives and policy parameters Execution Systems Field CRM · marketing automation digital platforms · partners FIELD_DETAIL · REP_EMAIL BROADCAST_EMAIL · THIRD_PARTY_MEDIA WEB_PORTAL · VIRTUAL_MEETING · EVENT Delivers approved action to the HCP Healthcare Professional Receives approved engagement and generates response signals Return signals Impression · open · click · dwell meeting · opt-out Commercial Analytics & Leadership Objectives · experiments guardrails · performance Sets optimisation policy and owns business outcomes Solid lines are operational flows. The dashed line is delayed engagement feedback — outcomes return hours to days after the action.
Who controls what
MLR & Governance
Controls what is permitted
Commercial Leadership
Controls what is optimised
Machine Learning Platform
Controls what is prioritised inside the eligible action set
Field Representative
Controls what is executed in the customer interaction
Why this matters
01

Reduces fragmented channel planning.

02

Improves allocation of field and marketing effort.

03

Creates a measurable feedback loop from action to incremental impact.

Key takeaway. The platform does not author content, override approval or replace the field team. It coordinates approved evidence, business goals and current HCP context into one auditable recommendation.

From operating model to decision contract

Now that the system boundary, ownership and feedback loop are clear, the next question is: what exactly is the unit of decision this platform produces?

View 02L0 · Decision

The six questions resolve into one recommendation

Six questions, one scored tuple. Building them as separate models is how a channel decision ends up contradicting a content decision.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'13px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart TB
  subgraph Q["Six connected questions"]
    direction LR
    q1["1 · Which HCP?
Next-Best-HCP
ranker + territory constraint"] q2["2 · What action?
Next-Best-Action
multi-objective ranker"] q3["3 · Which content?
Next-Best-Content
two-tower + graph retrieval"] q4["4 · Which channel?
Next-Best-Channel
contextual bandit"] q5["5 · When?
Next-Best-Time
survival / hazard model"] q6["6 · Why?
Explanation
graph path + feature attribution"] end q1 --> rec q2 --> rec q3 --> rec q4 --> rec q5 --> rec q6 --> rec rec["One recommendation record
hcp_id · action_type · content_id
channel · send_window · score
rationale · eligibility_proof"] rec --> crm["CRM surface
rep worklist"] classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.6px,color:#2E2470; classDef io2 fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.6px,color:#0A464C; class q1,q2,q3,q4,q5,q6 model; class rec,crm io2;

Question 6 is not an afterthought — eligibility_proof is emitted by the policy layer, not reconstructed later by an LLM.

View 03Data

Dataset, schema and feature construction

Scale is set by what each downstream method needs to be statistically meaningful: enough interactions per HCP for sequence modelling, enough randomisation units for cluster-randomised experiments, and enough treated/control contrast for uplift estimation.

Show full table — 13 rows by Entity
EntityScaleDrives
HCPs50,000User-side embeddings, segment strata
HCOs5,000Affiliation edges, institutional influence
Representatives500Territory capacity constraints
Territories100Cluster-RCT randomisation units
Brands / domains3Cross-domain transfer experiment
Channels5Bandit arm space
Content campaigns2,000Lineage root
Modules5,000Mid-lineage grouping
Content key messages10,000Item-side embeddings
Engagement events3–5 MInteraction matrix, sequences
Graph relationships5–10 MGraph retrieval, GNN training
Experiments20–30Uplift and policy evaluation
Time period18–24 monthsTemporal splits, seasonality, drift

Core event schema

TableFields
Eventevent_id · hcp_id · timestamp · session_id · brand_id · channel · action_type · content_id · topic_id · rep_id · territory_id · impression · open · click · dwell_seconds · follow_up_accepted · event_registered · opt_out
HCP featurehcp_id · specialty · segment · territory · affiliated_hco · historical_channel_affinity · topic_affinity_vector · engagement_frequency · days_since_last_contact · preferred_time_window · access_barrier_cluster · graph_embedding
Contentcontent_id · tactic_id · module_id · atom_id · topic · channel · brand · approval_status · effective_date · expiry_date · segment_eligibility · embedding
Leakage discipline. All splits are temporal, never random. Features are computed as of the event timestamp through point-in-time joins, and a dedicated leakage test asserts that no feature in a training row was observable only after that row's label.
View 04L2 · Graph

HCP-360 knowledge graph schema

Three retrieval paths the feature tables cannot serve: content lineage for asset-level attribution, affiliation for cold-start, and barrier matching for the highest-intent recommendation in the system.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  HCP(["HealthcareProfessional"])
  HCO(["HCO"])
  SEG(["HCPSegmentation"])
  REP(["Representative"])
  TER(["Territory"])
  ENG(["Engagement"])
  TAC(["Campaign"])
  MOD(["ContentModule"])
  ATOM(["KeyMessage"])
  BRAND(["Brand"])
  CH(["Channel"])
  TOP(["Topic"])
  BAR(["AccessBarrier"])
  EVT(["Event"])

  HCP -->|"HCP_AFFILIATED_WITH_HCO"| HCO
  HCP -->|"hasSegment"| SEG
  HCP -->|"HCP_SIMILAR_TO_HCP"| HCP
  HCP -->|"HCP_ATTENDED_EVENT"| EVT
  HCP -->|"HCP_FACES_BARRIER"| BAR
  ENG -->|"isDeliveredTo"| HCP
  TAC -->|"Delivers"| ENG
  TAC -->|"isComprisedOf"| MOD
  MOD -->|"isComprisedOf"| ATOM
  TAC -->|"promotes"| BRAND
  ENG -->|"hasChannel"| CH
  ATOM -->|"describes"| TOP
  ATOM -->|"CONTENT_ADDRESSES_BARRIER"| BAR
  REP -->|"REP_COVERS_TERRITORY"| TER
  HCP -->|"located in"| TER

  classDef semantic fill:#FAEDDD,stroke:#A85D08,stroke-width:1.5px,color:#6B3B05;
  classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470;
  class HCP,SEG,ENG,TAC,MOD,ATOM,BRAND,CH semantic;
  class HCO,REP,TER,TOP,BAR,EVT model;

Amber nodes are the core content and engagement ontology. Violet nodes support retrieval, territory planning and barrier-aware content matching.

Controlled vocabularies — taken verbatim from the ontology - 3 rows
EnumerationValuesRole in the recommender
omni:ChannelEnumFIELD_DETAIL · REP_EMAIL · BROADCAST_EMAIL · THIRD_PARTY_MEDIA · WEB_PORTAL · VIRTUAL_MEETING · EVENTAction space for the Next-Best-Channel bandit
omni:SegmentEnumEarlyAdopter · GrowthPotential · EstablishedEligibility filter and cold-start prior
omni:ApprovalEnumApproved · InReview · ExpiredHard policy gate — only Approved is servable

Graph retrieval methods compared

MethodPurposeCostExplainable
Rule traversalTransparent baseline; the thing every other method must beatLowFully — path is the reason
Personalised PageRankRank graph-near content and similar HCPsMediumYes — via path weight
Node2VecStructural HCP and content embeddingsMediumWeakly
GraphSAGELearns from node attributes plus neighbourhood; handles new nodesHighLimited — neighbour contribution analysis
LightGCNCollaborative signal propagated over the interaction graphHighWeakly
Heterogeneous graph transformerCross-entity recommendation over typed edgesHighestVia metapath attention
View 05L1 · Platform

End-to-end architecture

Two representations of one truth — a feature store for point lookups, a graph for traversal — feeding a single linear decision path.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  subgraph SRC["Sources"]
    direction TB
    s1["Content systems
Modular Content Authoring · MLR Workflow
Digital Asset Management · Campaign Taxonomy"] s2["Delivery systems
Field CRM (detail + broadcast)
Marketing Automation · Third-party media"] s3["Commercial warehouse
Commercial Data Warehouse"] s4["Consent & suppression
Consent Management · Suppression Service"] end subgraph LAKE["Lakehouse"] direction TB l1["Lake · Raw zone
append-only events"] l2["Lake · Refined zone
conformed entities"] l3["Spark / Delta
Hive-partitioned
point-in-time joins"] end subgraph FEAT["Feature & semantic layer"] direction TB f1["Offline feature store
time-aware training sets"] f2["Online store
Feast / Redis"] f3["HCP-360 knowledge graph
Graph DB · SPARQL"] f4["ANN index
FAISS / OpenSearch"] end subgraph CORE["Decision core"] direction TB c1["Candidate generation
5 retrievers"] c2["Learning-to-rank
LambdaMART"] c3["Sequence rescorer
SASRec / BERT4Rec"] c4["Bandit action selector
Thompson / LinUCB"] c5["Policy & eligibility filter
approval · consent · frequency"] end subgraph SERVE["Serving"] direction TB v1["Recommendation API
FastAPI"] v2["Pre-call copilot
Graph-RAG explanation"] v3["Control centre
dashboard"] end s1 -->|"~3K assets/qtr"| l1 s2 -->|"~500K events/day"| l1 s3 -->|"nightly, 2.1 TB"| l2 s4 -->|"consent deltas"| l2 l1 --> l3 l2 --> l3 l3 -->|"180M feature rows/night"| f1 l3 -->|"+40K edges/day"| f3 f1 -->|"500K HCP keys"| f2 f1 -->|"40K x 128-d"| f4 f3 --> c1 f4 --> c1 f2 -->|"340 features, 12ms"| c2 c1 -->|"~100 candidates"| c2 --> c3 --> c4 --> c5 c5 -->|"2,000 QPS · P95 180ms"| v1 c5 --> v2 c5 --> v3 v1 -->|"delivered action"| s2 s2 -.->|"engagement events - async, hours to days"| l1 s2 -.->|"real-time stream - 4.5K/s peak"| l3 classDef infra fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.5px,color:#0A464C; classDef semantic fill:#FAEDDD,stroke:#A85D08,stroke-width:1.5px,color:#6B3B05; classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; class s1,s2,s3,s4,l1,l2 infra; class f3 semantic; class l3,f1,f2,f4,c1,c2,c3,c4,c5,v1,v2,v3 model;

Teal = data and serving infrastructure · amber = the semantic layer · violet = learned decision components. The dotted return edge is asynchronous — engagement outcomes land hours to days later.

View 06L2 · Retrieval

Candidate generation — five retrievers, one union

Recall is a ceiling — ranking cannot recover what retrieval missed. Five retrievers with different failure modes emit channel-specific tuples; each is measured on its own Recall@50 so a silently broken retriever is visible rather than absorbed.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12.5px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  hcp["HCP request
hcp_id + context"] r1["Policy-aware baseline
coarse eligibility pre-filter
segment · brand · indication"] r2["Collaborative filtering
item-item on engagement
ALS / implicit"] r3["Two-tower ANN
HCP tower x content tower
FAISS top-200"] r4["Graph neighbours
personalised PageRank
metapath traversal"] r5["Segment popularity
cold-start fallback
EarlyAdopter/GrowthPotential/Established"] u["Union + dedupe
candidate_key = hcp_id + action_type
+ content_id + channel + send_window"] cap["Cap to 100 candidates
proportional per source"] hcp --> r1 hcp --> r2 hcp --> r3 hcp --> r4 hcp --> r5 r1 --> u r2 --> u r3 --> u r4 --> u r5 --> u u --> cap cap --> nxt["to ranking"] classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; classDef io fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.5px,color:#0A464C; class r1,r2,r3,r4,r5 model; class hcp,u,cap,nxt io;
RetrieverCoversFails onMetric
Policy-aware baselineCoarse eligibility pre-filter, full transparencyNo personalisation signalCoverage %
Collaborative filteringDense HCPs with long engagement historyCold-start HCPs and new contentRecall@50
Two-tower ANNSemantic match on content and HCP needNeeds retraining as content churnsRecall@100
Graph neighboursAffiliation, institution, similar-HCP pathsSparse subgraphs, orphan nodesCold-start Recall@50
Segment popularityNever returns empty — the floorPopularity bias, low diversityDiversity / coverage
View 07Serving · Lifecycle

Serving: request lifecycle and in-session adaptation

One request end to end, and how the features it reads are kept fresh between requests.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
sequenceDiagram
  autonumber
  participant CRM as Field CRM
  participant API as Recommendation API
  participant FS as Online store
  participant R as Retrieval tier
  participant M as Ranking models
  participant P as Policy gate
  CRM->>API: GET /recommendations?hcp_id
  API->>FS: fetch batch + session features
  FS-->>API: 340 features · 12 ms
  API->>R: fan out to 5 retrievers in parallel
  R-->>API: ~100 candidate tuples · 45 ms
  API->>M: score tuples
  M-->>API: LambdaMART 100 · 18 ms
  M-->>API: sequence rescore top 20 · 55 ms
  API->>API: utility adjust + bandit select · 5 ms
  API->>P: evaluate eligibility
  P-->>API: pass or suppress + proof · 9 ms
  API-->>CRM: top K + rationale · P95 180 ms

Latency budget — allocated targets per stage

Full latency budget - 11 rows
StageP50P95Notes
Request parse and auth1 ms2 ms
Feature fetch — online store4 ms12 msBatch + session features in one round trip
Candidate generation — 5 parallel14 ms45 msBounded by graph traversal, the slowest retriever
Union, dedupe, proportional cap1 ms3 ms
LambdaMART — 100 candidates7 ms18 msScales linearly in candidate count
Sequence rescorer — top 20 only12 ms55 msDominant cost; bounded by rescoring the head, not all 100
Utility adjustment<1 ms1 msArithmetic on business coefficients
Bandit arm selection2 ms4 msPosterior sample over 7 arms
Policy gate — 6 rules3 ms9 msShort-circuit; most candidates fail rule 1
Response assembly and rationale2 ms5 ms
Total45 ms180 msAgainst a 250 ms budget

In-session adaptation - how the features that request reads stay current

The streaming path changes exactly one thing: which features the ranker reads at request time. It retrains nothing.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12.5px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  ev["CRM & web events
open · click · dwell · page view"] k["Kafka
topic: hcp.engagement.v1"] ss["Spark Structured Streaming
sessionise by hcp_id
10-min watermark"] fu["Session feature updater
recent topics · channel · time gaps"] on["Online store
Feast / Redis"] api["Rescoring API"] hist["Delta history
replay + training"] ev --> k --> ss --> fu --> on --> api ss --> hist api -->|"refreshed ranked list"| ui["Rep worklist / web surface"] classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; classDef io fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.5px,color:#0A464C; class k,ss,fu,on,api model; class ev,hist,ui io;
Worked example. Initial recommendation: general educational content. New event arrives — the HCP opens an access-related page and searches reimbursement terms. Session features update within the watermark; the next request retrieves access-resource content modules and a field-access follow-up rises to rank 1.

Failure modes and degradation

Component failsDetected byBehaviourUser impact
Session feature streamConsumer lag alert > 60 sAPI proceeds on batch features onlySlightly less responsive to in-session intent
ANN index unavailablePer-source Recall@50 drops to zeroFour remaining retrievers serveReduced semantic coverage, list still full
Graph database timeoutRetriever P95 breachRetriever drops out at 60 ms deadlineLoses relational and cold-start paths
Sequence rescorerStage latency breachFalls back to LambdaMART orderingLoses short-term intent signal
Bandit serviceHealth checkFalls back to highest-prior channelExploration pauses, exploitation continues
Policy gateHealth checkRequest fails closed — no recommendation servedEmpty worklist
Online storeRead error rateReuse last-known-good scores, then re-run the authoritative gate against current approval, consent and contact state. If policy state is also unavailable, fail closedOrdering may be stale; eligibility never is

Ranking quality degrades gracefully; eligibility does not degrade at all. A cached score is acceptable, a cached eligibility decision is not — content expires, consent is withdrawn and frequency limits are reached between requests.

View 08L2 · Ranking

Ranking and the policy cascade

Retrieval emits channel-specific tuples, so channel is scored rather than assigned afterwards. Exploration enters as a bonus term on an already-eligible tuple — it never reassigns a component of a scored candidate. The authoritative eligibility gate runs last and is absolute.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12.5px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart TB
  cands["~100 channel-specific tuples
from retrieval"] --> ltr ltr["Stage 1 · LightGBM LambdaMART
ranks full tuples
HCP x action x content x channel
meaningful-engagement relevance labels
objective: NDCG@K"] seq["Stage 2 · Sequence rescorer
SASRec over recent interaction sequence
blends short-term intent"] util["Stage 3 · Utility and exploration
score + exploration bonus
− fatigue − opt-out risk − cost"] ban["Stage 4 · Final ordering
argmax over eligible tuples
channel is carried, never reassigned"] ltr --> seq --> util --> ban --> pol subgraph POL["Stage 5 · Policy gate — absolute, non-overridable"] direction TB p1["approval_status = Approved"] p2["effective_date <= today < expiry_date"] p3["segment_eligibility contains HCP segment"] p4["channel permitted for content type"] p5["consent present · not opted out"] p6["contact frequency within limit"] end pol["eligibility evaluation"] --> p1 --> p2 --> p3 --> p4 --> p5 --> p6 p6 --> out["Final ranked list
top K + suppression log"] classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; classDef policy fill:#FAEDDD,stroke:#A85D08,stroke-width:1.6px,color:#6B3B05; classDef io fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.5px,color:#0A464C; class ltr,seq,util,ban model; class p1,p2,p3,p4,p5,p6,pol policy; class cands,out io;

Every suppression records the full candidate tuple, the rule violated, the policy version and the evaluation timestamp — {recommendation_id, hcp_id, action_type, content_id, channel, send_window, rule_violated, policy_version, evaluated_at}. Suppression compliance is a first-class metric, not an exception path.

View 09L2 · Validation

Validation strategy — four tracks, each against its own baseline

These components solve different problems and are not rungs on one ladder. A retriever cannot beat a ranker: one is measured on recall, the other on ordering. Each track has its own baseline, its own metric, and its own promotion rule.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  subgraph T1["Retrieval track - Recall@50/100, cold-start recall, coverage"]
    direction LR
    a0["Segment
popularity"] --> a1["Collaborative
filtering"] --> a2["Two-tower
retrieval"] --> a3["Graph-hybrid
retrieval"] end subgraph T2["Ranking track - NDCG@10, MAP@K, calibration"] direction LR b0["Rules
score"] --> b1["Propensity
model"] --> b2["LambdaMART"] --> b3["Sequence
rescoring"] end subgraph T3["Online decision track - cumulative reward, regret, opt-out rate"] direction LR c0["Fixed channel
policy"] --> c1["Epsilon-greedy"] --> c2["LinUCB"] --> c3["Thompson
sampling"] end subgraph T4["Causal track - Qini, uplift@K, incremental per 1K HCPs"] direction LR d0["Propensity
targeting"] --> d1["T-learner"] --> d2["X-learner"] --> d3["Causal
forest"] end classDef io fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.5px,color:#0A464C; classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; class a0,b0,c0,d0 io; class a1,a2,a3,b1,b2,b3,c1,c2,c3,d1,d2,d3 model;

Teal is the incumbent baseline for each track — what the business does today, or the simplest thing that works. Promotion requires beating that baseline on the track's own metric, on the same temporal split, with bootstrap confidence intervals over HCPs.

TrackQuestion it answersBaseline to beatPrimary metricPromotion rule
RetrievalDoes the right action reach the ranker at all?Segment popularityRecall@100 · cold-start Recall@50Raises the recall ceiling enough to justify an index rebuild
RankingIs the ordering right?Rules scoreNDCG@10 · MAP@KBeats current practice, and calibration does not degrade
Online decisionWhich channel, under uncertainty?Fixed channel policyCumulative reward · regretLower regret without raising opt-out rate
CausalDid the action cause the outcome?Propensity targetingQini · uplift@KHigher incremental engagement per 1,000 HCPs
The rule is "beat the right baseline", not "beat the previous model". Scoped adoption is a valid outcome — if two-tower retrieval wins only on cold-start HCPs, it serves cold-start candidates and nothing more. What does not happen is a global claim from a segment win.
View 10L2 · Online

Online learning — exploration inside the eligible set

Offline learning can only learn from actions that were taken, which leaves channel preference confounded by historical policy. Exploration is the only mechanism that breaks that — and it runs strictly inside the eligible action space.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12.5px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  ctx["Context
HCP features · current sequence
channel · time · territory
content category · contact frequency"] pol["Policy
Thompson sampling
posterior over arm reward"] arms["Arms
email content A · email content B
rep follow-up · event invitation
web resource · no action"] obs["Observed outcome"] rew["Reward"] upd["Posterior update"] ctx --> pol --> arms --> obs --> rew --> upd --> pol classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; classDef io fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.5px,color:#0A464C; class pol,arms,upd model; class ctx,obs,rew io;

Reward function

OutcomeRewardWhy weighted this way
Email open+0.1Weak signal, easily inflated by preview panes
Click+0.3Intent, but not yet value
Meaningful content engagement+0.6Dwell past threshold — the real objective
Follow-up accepted+1.0Terminal success for the action
Opt-out−1.0Loses the channel permanently
No action taken0.0A real arm with a real outcome — restraint is learnable
Exploration is bounded to the eligible set. Arms that would breach approval, consent, channel permission or contact frequency are removed before the policy samples, so a compliance breach is not a low-reward outcome the bandit learns to avoid — it is unreachable. What the policy does learn from is engagement, dwell, follow-up acceptance, fatigue and opt-out.

Baselines compared: fixed channel policy · epsilon-greedy · LinUCB · Thompson sampling. Reported on cumulative reward, regret, action diversity and opt-out rate.

View 11L2 · Causal

Experimentation and uplift

A propensity model asks who is likely to respond. An uplift model asks who responds because we acted. Targeting the first produces a system that takes credit for engagement that would have happened anyway — and, worse, keeps contacting sleeping dogs.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12.5px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart TB
  subgraph DES["Experiment design"]
    direction LR
    d1["HCP-level randomisation
when no rep spillover"] d2["Cluster RCT by territory
when reps serve many HCPs"] d3["Switchback by time block
when policy is global"] end DES --> run["Run: personalised policy
vs rules / standard segmentation"] run --> outc["Outcomes
engagement · acceptance
fatigue · opt-out rate"] outc --> up["Uplift model
T-learner · X-learner · causal forest"] up --> g1["Persuadables
respond only if contacted
→ TARGET THESE"] up --> g2["Sure things
respond regardless
→ do not spend contact"] up --> g3["Lost causes
never respond
→ do not spend contact"] up --> g4["Sleeping dogs
contact harms engagement
→ actively suppress"] classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; classDef infra fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.8px,color:#0A464C; classDef policy fill:#FAEDDD,stroke:#A85D08,stroke-width:1.6px,color:#6B3B05; class d1,d2,d3,run,outc,up model; class g1 infra; class g4 policy; class g2,g3 model;

Reported on Qini curve, uplift@K, and incremental engagement per 1,000 HCPs — never on raw response rate in the treated group.

View 12L2 · GenAI

Compliant pre-call copilot — where the LLM is and is not allowed

This is the single most important boundary in the architecture. The LLM does not rank, does not select, and does not decide eligibility. It receives an already-ranked, already-filtered list plus the approved evidence, and writes the explanation. Saying "I used an LLM to recommend content" describes a system that cannot pass MLR review.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12.5px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart TB
  q["Rep question
'best approved follow-up for this HCP
given recent access-content activity?'"] subgraph DET["Deterministic zone — ML and rules decide"] direction TB a1["1 · Retrieve HCP history
graph traversal"] a2["2 · Generate candidates
five retrievers"] a3["3 · Rank
LTR + sequence rescorer"] a4["4 · Filter
approval · consent · recency · frequency"] a5["5 · Retrieve supporting evidence
hybrid vector + metadata + graph"] end subgraph GEN["Generative zone — language only"] direction TB b1["6 · Explain the ranked list
in natural language"] b2["7 · Cite key message IDs and
the feature signals used"] end q --> a1 --> a2 --> a3 --> a4 --> a5 --> b1 --> b2 --> out["Answer with citations"] note["Boundary rule
the LLM never reorders, never adds a candidate,
never overrides an eligibility decision"] GEN -.- note classDef infra fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.6px,color:#0A464C; classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.6px,color:#2E2470; classDef policy fill:#FAEDDD,stroke:#A85D08,stroke-width:1.8px,color:#6B3B05; class a1,a2,a3,a4,a5 infra; class b1,b2,out model; class note policy;

Evaluated on factuality, citation correctness, groundedness, constraint satisfaction, and hallucination rate — not on fluency.

View 13L2 · MLOps

Observability

A recommender degrades quietly. Ranking quality decays before anyone reports a problem, and subgroup damage — a specialty or a segment being systematically under-served — never shows up in an aggregate metric.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12.5px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  subgraph SIG["Signals"]
    direction TB
    x1["Feature drift
PSI per feature"] x2["Prediction drift
score distribution"] x3["NDCG decay
rolling window"] x4["Training/serving skew"] x5["Subgroup performance
by specialty · segment · territory"] x6["Latency P50 / P95"] x7["Suppression rate
by rule"] end reg["Model registry
MLflow"] mon["Monitoring service"] alert["Alerting"] act["Actions
retrain · rollback · freeze policy"] SIG --> mon --> alert --> act reg --> mon act --> reg classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; classDef io fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.5px,color:#0A464C; class x1,x2,x3,x4,x5,x6,x7 model; class reg,mon,alert,act io;
View 14Close

What this demonstrates, what it costs, and what comes next

What the design establishes, the costs it accepts to get there, and the three things I would do first with production traffic.

What the system demonstrates

CapabilityHow it shows up in the design
A coherent joint decisionOne scored tuple - HCP, action, content, channel, window - rather than components reconciled after the fact
Hybrid retrievalFive retrievers with complementary failure modes, each measured on unique contribution
Ranking under business utilityModel score adjusted by commercially-owned fatigue, opt-out and cost coefficients held outside the learned objective
Deterministic complianceA zero-error eligibility gate that runs last, emits an audit proof, and fails closed
Bounded explorationOnline learning confined to the eligible action space, so a breach is unreachable rather than penalised
Real-time adaptationSession features refreshed within minutes on a path that degrades gracefully to batch
Causal measurementTargeting on incremental rather than predicted response, with randomisation designed around field spillover
Grounded explanationGeneration scoped to language, with citations validated against the retrieved evidence set

Accepted trade-offs

DecisionWhat it costsWhy it was still the right call
Joint decision recordPipeline coupling — the channel model cannot be released independently of the rankerIndependent models produce contradictory recommendations, and the reconciliation layer becomes unlearnable
Penalties outside the learned objectiveSome loss of joint optimality against a fully-learned objectiveContact policy is a commercial accountability; the business can retune without a model release
Eligibility gate lastScoring work spent on candidates that are later suppressedFreshest possible approval check, an auditable suppression log, and no compliance signal leaking into the ranking objective
Two stores over oneConsistency risk between the feature store and the graphPoint lookup and variable-depth traversal are different access patterns; one store serves one of them badly
Generative layer scoped to languageLoses whatever contextual signal a language model could contribute to orderingReproducibility is required for audit and for experiment validity; that signal is better distilled into an offline feature
RDF over a property graphTraversal throughput I have not benchmarked at production scaleFormal ontology governance and interoperability across many source systems

What I would change first

#ChangeReason
1Basic observability from first deployment — logging, lineage, model versioning and latency monitoring ship with the first served model; drift and subgroup analysis follow once production outcomes matureLightweight logging and latency monitoring should exist from the first served model; only drift and subgroup analysis need a running system to be meaningful
2Benchmark the graph store against a property-graph alternativeRetrieval latency at production traversal volume is an open question, not a settled one
3Re-engineer the ANN index and graph rebuilds to be incrementalBoth are full recomputations today and are the first components that would break at ten times the volume
4Calibrate reward and penalty coefficients against realised opt-out ratesThe current values encode a defensible ordering but are set a priori rather than fitted
5Add a learned retrieval blend in place of proportional allocationProportional allocation is transparent but almost certainly leaves recall on the table

Open questions I would answer with production data

  • Does the shared cross-channel encoder actually beat per-channel models on sparse-channel HCPs, or is the aggregate gain driven entirely by the dense channels?
  • How large is the negative-uplift population in practice — is suppressing it a material effect or a rounding error?
  • Does in-session feature refresh change decisions often enough to justify the streaming path over hourly batch?
On measurement. Offline ranking quality and business impact are different questions and can disagree. Offline NDCG measures ordering against held-out engagements under the existing logging policy; whether that translates into incremental value is answered only by randomised experiment, reported as Qini and incremental engagement per thousand HCPs.

Appendix

Reference detail - capacity tables, transfer architecture, repository layout and build sequencing. Available for questions rather than presented in flow.

Appendix 1Capacity

Scale, throughput and service objectives

Design capacity and service targets, not recorded benchmarks — these are the figures the system is sized and engineered against. Batch, streaming and request-time paths carry very different load profiles, and each has its own objective — throughput for batch, freshness for streaming, latency for serving.

Entity and corpus volume

EntityVolumeGrowthWhat it drives
Healthcare professionals500 K~3% / yrUser-side embeddings, segment strata
Healthcare organisations45 K~2% / yrAffiliation edges, institutional influence
Field representatives4,500flatTerritory capacity constraint
Territories900realigned 2× / yrCluster-randomisation units
Campaigns8 K~600 / quarterLineage root
Content modules20 K~1.5 K / quarterMid-lineage grouping
Key messages40 K~3 K / quarterItem-side embeddings, retrieval corpus
Engagement events180 M / yr~500 K / dayInteraction matrix, sequences, labels
Graph nodes / edges1.2 M / 14 M+40 K edges / dayGraph retrieval, GNN training
Lakehouse footprint2.1 TB~55 GB / monthRaw + refined + feature tables

The three load profiles

PathVolume & shapeComputeObjectiveSLO
Batch
nightly training + features
180 M feature rows / night
2.1 TB scanned
Spark on Hive-partitioned Delta tables · 40 executors · 8 vCPU / 32 GBThroughput and point-in-time correctnessTarget: complete in < 45 min
Streaming
engagement events
~500 K events / day sustained
peak 4.5 K events/sec during campaign sends
Kafka 24 partitions, keyed by hcp_id · Spark Structured Streaming, 10-min watermarkFeature freshnessConsumer lag < 60 s
event→feature < 5 min
Session
in-flight interaction state
~85 K concurrent sessions
~4 KB state each · 340 MB working set
Redis, 30-min TTL, keyed by hcp_idRead latencyP99 read < 5 ms
Serving
recommendation API
2,000 QPS peak
~200 K candidate scorings / sec
FastAPI, horizontally scaled · warm model cacheLatency and availabilityTarget P50 45 ms · P95 180 ms
Target availability 99.9%

The event stream is bursty rather than uniform: a broadcast send to a large segment produces hundreds of thousands of impression events in minutes, which is why the streaming tier is sized on peak rather than mean.

Model and index operations

ArtefactSizeCadenceBuild cost
Ranker training set~85 M labelled rows · 340 featuresWeekly~35 min
Two-tower embeddings500 K HCP × 128-d · 40 K item × 128-dNightly~50 min GPU
ANN index40 K vectors · ~40 MBNightly full rebuild~4 min
Sequence modelMax sequence 200 · 6 layersWeekly~2.5 h GPU
Graph rebuild1.2 M nodes · 14 M edgesNightly, edges incremental~18 min
Bandit posteriors7 arms × context dimContinuousstreaming update
Known scaling limit. The ANN index and graph rebuilds are full recomputations. At roughly ten times current volume both become the binding constraint on nightly window and would need incremental construction — this is the first thing I would re-engineer.
Appendix 2Transfer

Cross-brand and cross-channel transfer

The claim to test: behaviour in one domain improves prediction in another, most visibly for HCPs with thin history in the target domain. The ablation is the experiment — a shared encoder that does not beat per-channel models on sparse users has not earned its complexity.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12.5px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  subgraph IN["Per-HCP input"]
    direction TB
    i1["Profile
specialty · segment · territory"] i2["Interaction sequence
all brands, all channels"] i3["Graph embedding
Node2Vec / GraphSAGE"] end enc["Shared HCP encoder
transformer over unified sequence"] h1["Email head
BROADCAST_EMAIL / REP_EMAIL"] h2["Field-action head
FIELD_DETAIL rep call"] h3["Web-content head"] h4["Event head
THIRD_PARTY_MEDIA / educational"] i1 --> enc i2 --> enc i3 --> enc enc --> h1 enc --> h2 enc --> h3 enc --> h4 abl["Ablation ladder
channel-only → brand-only
→ shared → shared+graph
→ shared+session transformer"] enc -.-> abl classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.5px,color:#2E2470; classDef io fill:#E1F0F1,stroke:#0E7C86,stroke-width:1.5px,color:#0A464C; class enc,h1,h2,h3,h4 model; class i1,i2,i3,abl io;

Three commercial domains: Brand A clinical education · Brand B access and reimbursement · Brand C scientific events.

Appendix 3Code layout

Repository map

One repository, connected content modules — not twelve disconnected notebooks. Each folder maps to a component in View 03 and each is importable by the ones downstream of it.

%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace,SFMono-Regular,Menlo,monospace','fontSize':'12px','lineColor':'#5C6675','primaryTextColor':'#0D1117'}}}%%
flowchart LR
  d01["01_data_generation"] --> d02["02_spark_feature_pipeline"]
  d02 --> d03["03_behavioral_analysis"]
  d02 --> d04["04_baseline_recommenders"]
  d02 --> d05["05_candidate_generation
collaborative_filtering · two_tower · graph_retrieval"] d05 --> d06["06_learning_to_rank"] d06 --> d07["07_sequential_transformer"] d07 --> d08["08_cross_domain_model"] d02 --> d09["09_streaming_personalization"] d06 --> d10["10_contextual_bandit"] d10 --> d11["11_experimentation"] d05 --> d12["12_rag_copilot"] d06 --> d13["13_serving_api"] d08 --> d13 d09 --> d13 d13 --> d14["14_monitoring"] d13 --> d15["dashboard"] d11 --> d15 classDef model fill:#EBE8FA,stroke:#5B4BC4,stroke-width:1.4px,color:#2E2470; class d01,d02,d03,d04,d05,d06,d07,d08,d09,d10,d11,d12,d13,d14,d15 model;

Folder to component mapping

Show full table — 15 rows by Folder
FolderArchitecture componentView
01_data_generationHCP, content, event and graph data generation15
02_spark_feature_pipelineLakehouse + offline feature store on Hive-partitioned Delta, point-in-time joins03
03_behavioral_analysisJourney analytics, Markov chains, cohort analysis03
04_baseline_recommendersRules, popularity, CF, matrix factorisation08
05_candidate_generationFive retrievers + union + ANN index04
06_learning_to_rankLambdaMART stage-1 ranker05
07_sequential_transformerSASRec / BERT4Rec sequence rescorer05
08_cross_domain_modelShared HCP encoder + per-channel heads09
09_streaming_personalizationKafka → Spark SS → Feast/Redis06
10_contextual_banditThompson / LinUCB channel selection10
11_experimentationA/B framework, switchback, uplift models11
12_rag_copilotGraph-RAG explanation layer12
13_serving_apiFastAPI + policy gate + online store05
14_monitoringDrift, NDCG decay, subgroup performance13
dashboardPersonalization control centre13
Appendix 4Delivery plan

Build order

Ordered so that every phase produces something demonstrable and nothing is blocked on the phase after it.

Full build sequence - 12 rows
PhaseBuildComponentDemonstrates
1Data generation, temporal splits, EDAData & feature pipelineData mining, user-behaviour modelling
2Popularity, CF, matrix-factorisation baselinesRetrieval & ranking baselinesRecommendation fundamentals
3Multi-source candidate generator, ANN indexCandidate generationRetrieval and candidate generation
4LightGBM learning-to-rank + logging, lineage, model versioning and latency monitoringRankingRanking end-to-end
5Knowledge graph and graph retrievalKnowledge graphGraph-based recommendation
6SASRec / BERT4Rec sequence modelSequence rescoringTransformers and attention
7Shared cross-brand / cross-channel modelCross-channel modelCross-vertical personalization
8Kafka, Spark Streaming, Redis featuresReal-time adaptationReal-time session adaptation
9Contextual bandit and uplift modellingOnline learning & causalOnline and incremental learning
10A/B framework and monitoringExperimentation & full drift/subgroup monitoringExperimentation and product judgement
11Grounded RAG pre-call copilotExplanation layerLLMs in personalization, done compliantly
12FastAPI, dashboard, this documentServing & observabilityProduction readiness