How a top-five public university stopped guessing what its alumni cared about — and started reading it directly from their clicks.

By Pavan Malpani and Sandeep Aradada

SUMMARY:

  • Open rates and click counts measure activity. They do not tell you what the activity meant.
  • Classifying every campaign link against a versioned taxonomy before the send turns each click into a topic signal.
  • At a top-five public university: 26,000 links classified against a 180-node taxonomy, at 94% precision.
  • Interest-targeted sends produced a 41% higher click-through rate, and building an audience fell from six hours to under twenty minutes.
  • Consent, score decay and per-signal attribution are part of the design, not added afterwards.

Table of Contents

  1. Why Open Rates Are Not Audience Intelligence
  2. What Audience Intelligence Looked Like at a Top-Five Public University
  3. How a Campaign Click Becomes an Audience Intelligence Signal
  4. What the Audience Intelligence Output Actually Looks Like
  5. What Audience Intelligence Unlocks, by Sector
  6. Privacy, Consent, and Governance
  7. What It Costs and How Accurate It Is
  8. Frequently Asked Questions
  9. The Bottom Line

Why Open Rates Are Not Audience Intelligence

Organizations running mass email campaigns — from university alumni offices to enterprise marketing teams — all hit the same wall after send: they know who opened and who clicked, but not what those clicks actually mean. Open rates and click totals are activity metrics. They do not explain interest.

A click stays an ambiguous signal until the link behind it is classified against a structured topic taxonomy. Once every outbound link carries a versioned, standardized tag, clicks stop being noise and start becoming data — individual signals that, aggregated per recipient, build a genuine interest graph you can segment and personalize against.

What Audience Intelligence Looked Like at a Top-Five Public University

XTIVIA built and deployed this for the advancement office of one of the five largest public universities in the United States — an institution with an alumni base of 700k+ and a campaign calendar of roughly 3,500+ sends a year across colleges, athletics, and central development. The work sits inside our broader enterprise AI solutions practice.

Before

Link reporting stopped at the URL. Staff could see that four thousand people clicked something, then read URLs in a spreadsheet to guess what it meant. Segmentation ran on giving history and class year — proxies for interest, not evidence of it. Building the audience for a single themed appeal took about six staff-hours of manual list work, and the result was really a guess dressed up as a segment.

After

26,000 unique campaign links classified against a 180-node taxonomy. 94% tagging precision measured against a human-reviewed sample. A per-recipient interest profile refreshed nightly. Interest-targeted sends produced a 41% higher click-through rate than the control segment, and the six hours of list work fell to under twenty minutes.

The short version. Nothing about the campaigns changed. The same emails went to the same people. What changed is that the organization could finally read back what the clicks meant — and act on it before the next send.

How a Campaign Click Becomes an Audience Intelligence Signal

Follow a single link. An alumni newsletter goes out with fourteen links in it. One points to a feature on the engineering school’s robotics lab. Three thousand people click it. In most advancement shops, that click arrives in a report as a URL and a number, and that is where it stops.

In this framework, that URL was crawled, stripped to its body text, and classified before the send ever went out — tagged Research › Engineering › Robotics and Student Life › Undergraduate Research, against taxonomy version 14. So the moment those three thousand clicks land, three thousand alumni records each gain a weighted interest signal rather than a URL. Do that across fourteen links and more than three and a half thousand sends a year, and you stop guessing what your audience cares about.

The classification work is where the engineering lives — headless rendering for JavaScript-heavy pages, boilerplate stripping, batched inference against a versioned taxonomy, and a state layer that lets any stage fail and retry without stalling the rest. If that is the part you care about, it is the subject of the companion article: Inside the Auto-Tagging Pipeline: Classifying Campaign Links at Enterprise Scale. Everything below is about what you do with the output.

What the Audience Intelligence Output Actually Looks Like

The end product is not a dashboard. It is a matrix: one row per recipient, one column per topic, one affinity score per cell, rebuilt on a schedule. Every segment anyone draws — for an appeal, an event invitation, a suppression list — is a query against it.

Two properties matter more than they look. Scores are weighted by recency and decay on a configurable window, so a profile reflects current interest rather than a permanent behavioral record. And every score is attributable — it traces back to a specific link, a specific taxonomy version, and a specific model. That second property is what makes the governance section below possible at all.

Closing the loop: getting segments back into the sending platform

A matrix sitting in a lakehouse changes nothing on its own. Computed segments are written back into the sending platform on a scheduled sync — Salesforce Marketing Cloud, Emma — so the next campaign’s audience selection and content blocks are driven by demonstrated interest rather than static list membership.

Suppression works the same way in reverse. Recipients with no affinity for a topic stop receiving that variant, which protects list health as much as it lifts engagement. In practice that second effect is the one advancement directors notice first — fewer irrelevant sends per person, not more.

What Audience Intelligence Unlocks, by Sector

  • Higher education. Surface alumni affinity toward specific research areas, athletic programs, or endowment initiatives — and route those affinities to the right gift officer before the next appeal goes out, including inviting alumni to in-person events based on affinity score for maximum participation.
  • Non-profits. Distinguish supporters drawn to policy advocacy from those drawn to direct volunteer work, and stop sending both groups the same appeal.
  • Commercial enterprise. Build product-affinity segments from real content engagement rather than from form fills and self-reported interest, which decay faster and lie more.

Building per-recipient interest profiles from behavioral data is a governance question before it is an engineering one — especially in higher education and non-profits, where the entire relationship runs on trust. Four design decisions carry most of the weight:

  • Consent is read at segmentation time, not at send time. A recipient who opts out of behavioral targeting is excluded from interest-based selection entirely, rather than filtered out at the last step.
  • Interest signals decay. Affinity scores age out on a configurable window, so a profile reflects current interest rather than a permanent behavioral record.
  • Every signal is attributable. Each entry traces back to a specific link, taxonomy version, and model — which is what makes a subject access request answerable in minutes instead of weeks.
  • The model never sees a person. Classification evaluates page content, not recipients. No recipient identifier is included in any model call; the join to individuals happens afterward, inside the client’s own tenant.

What It Costs and How Accurate It Is

Two questions come up in the first ten minutes of every conversation about this, so here they are without the hedging.

Accuracy. Every classification returns a confidence score. Assignments above 0.80 write straight through; anything below routes to a review queue where a staffer confirms or corrects the tag, and those corrections become the validation set for the next taxonomy revision. At the university, 87% of links cleared the auto-accept threshold, and a human-reviewed sample of 500 pages measured 94% precision. Accuracy depends far more on taxonomy quality than on model choice — a vague taxonomy produces vague tags no matter which model you point at it.

Cost. Roughly $0.004 per link at steady state — under $20 a month at this volume, against a 30-day re-crawl cadence for content that changes. Batching, boilerplate stripping, and caching against the link account for most of the gap between that and a naive implementation, which typically runs several times higher. The engineering behind those three levers is covered in the companion technical article.

Frequently Asked Questions

How is this different from the content tagging my email platform already does?

Most platforms tag at the campaign or link level using labels a human typed in when the campaign was built. That works until you have more links than people willing to label them, and it captures what the sender meant rather than what the content says. This classifies the destination page itself against a governed taxonomy, with versioning and confidence scoring — so tagging scales past the point where manual labeling stops, and the signals are consistent enough to aggregate per person.

Do we need a data lakehouse before we can start?

Not to begin. The classification pipeline can run against campaign exports and write to whatever store you already have. The lakehouse matters when you want the interest matrix joined to giving history, event attendance, and CRM data — which is where most of the value ends up, but it is not the first step.

How long does implementation take?

About four months to a working pipeline against a first taxonomy, and roughly three more campaign cycles before profiles are dense enough to segment on. You need several sends of click data before the graph says anything useful. Teams with a lakehouse already in place start from the second stage.

Is behavioral interest profiling GDPR- and CCPA-compliant?

The architecture is designed to support compliance, but compliance is a function of your consent capture, retention policy, and privacy notice — not of the pipeline alone. What the framework supplies is the part usually missing: per-signal attribution, configurable decay, opt-out enforcement at segmentation time, and no recipient identifiers in any model call. The obligations themselves are set out by the European Commission for GDPR and the California Attorney General for CCPA.

Can this run on our platform rather than AWS?

Yes. The university implementation runs on an AWS data lake, but the pattern is four decoupled stages over a medallion lakehouse and ports to Microsoft Fabric or Databricks without redesign. The stage-by-stage mapping is in the companion technical article.

The Bottom Line

Open rates and click counts tell you that engagement happened. Link tagging tells you what it meant — turning raw campaign activity into structured, per-recipient audience intelligence you can actually build a segmentation strategy on.

Ready to see what your campaign data is really telling you? Send us 100 links from your last campaign. We will classify them against a draft taxonomy and return a sample interest matrix — no engagement required to find out what your click data has been saying all along. Get in touch.

Keep reading — Part 2: Inside the Auto-Tagging Pipeline: Classifying Campaign Links at Enterprise Scale. The architecture behind the interest graph: four decoupled stages, a versioned taxonomy, strict output contracts, and how the same pattern runs on AWS, Microsoft Fabric, or Databricks.

Or explore our related service offerings: Enterprise AI Solutions, Microsoft Fabric Consulting, and Databricks Consulting.