← Back to Projects
Selected case study · Research data products

Custom Audiences

Custom Audiences are a product designed to predict likely audience membership across survey data and turn it into repeatable audience definitions.

FocusAudience modeling and productization
Result5,000+ audiences modeled for Fortune 500 clients; 23% of Upsell MRR in H1 2025
Problem

Anonymous panelists made cross-survey audience building a prediction problem.

The same person could not be directly joined from one survey to another. The system instead learned from an audience label in one survey and predicted likely members in another using demographic and cultural variables shared across datasets.

Solution

One system for modeling, profiling, and recurring production.

Designed an internal AutoML pipeline Python package for reusable model training and evaluation across classification and similarity-based lookalike workflows. Built an internal Python package that uses trained models to create audiences in the platform, including the necessary data transformation and ETL workflow. Helped productize recurring audience production and deployment, with sizing that applies survey incidence to ACS population data and BLS Consumer Expenditure estimates adding behavioral context where useful.

01

Classification modeling

An internally designed AutoML pipeline Python package supports reusable model selection, validation, evaluation, and experiment tracking for audience labels across survey data.

02

Similarity and lookalikes

A separate path ranks likely lookalikes using distance and overlap when a clean classification target is not the right fit.

03

Profiles and deployment

An internal Python package uses trained models to create structured, validated, platform-ready audiences, including the data transformation and ETL workflow needed to carry model output into the platform.

04

Sizing and context

Survey incidence and ACS population data support sizing, while BLS Consumer Expenditure estimates add behavioral context.

Reusable model evaluation

The internal AutoML pipeline package handles feature preparation, model selection, validation, evaluation, and experiment tracking so a method can be reused across audience definitions.

Classification and similarity

The classification path predicts likely audience membership from shared variables. The similarity path ranks lookalikes using distance and overlap when a clean target class is not available. Both workflows use a proprietary scoring and recommendation function: machine-learning evaluation metrics inform a custom score with penalties, helping select models and rank lookalikes for cross-survey use.

From model to audience product

The internal audience-creation package carries the chosen method into a structured, validated, platform-ready profile, including the data transformation and ETL workflow needed to create audiences in the platform. Recurring production and deployment keeps those audience definitions current, while sizing combines survey incidence with ACS population data. The customized Consumer Expenditure ETL transforms BLS datasets into readable tabular expenditure data, so audience predictions can be applied on top.

Value

Reusable audience definitions delivered at client scale.

The system turned questions about audience traits, media behaviors, and locations into reusable client audience definitions. More than 5,000 audiences were modeled for Fortune 500 clients, contributing 23% of Upsell MRR in H1 2025.