Classification modeling
An internally designed AutoML pipeline Python package supports reusable model selection, validation, evaluation, and experiment tracking for audience labels across survey data.
Custom Audiences are a product designed to predict likely audience membership across survey data and turn it into repeatable audience definitions.
The same person could not be directly joined from one survey to another. The system instead learned from an audience label in one survey and predicted likely members in another using demographic and cultural variables shared across datasets.
Designed an internal AutoML pipeline Python package for reusable model training and evaluation across classification and similarity-based lookalike workflows. Built an internal Python package that uses trained models to create audiences in the platform, including the necessary data transformation and ETL workflow. Helped productize recurring audience production and deployment, with sizing that applies survey incidence to ACS population data and BLS Consumer Expenditure estimates adding behavioral context where useful.
An internally designed AutoML pipeline Python package supports reusable model selection, validation, evaluation, and experiment tracking for audience labels across survey data.
A separate path ranks likely lookalikes using distance and overlap when a clean classification target is not the right fit.
An internal Python package uses trained models to create structured, validated, platform-ready audiences, including the data transformation and ETL workflow needed to carry model output into the platform.
Survey incidence and ACS population data support sizing, while BLS Consumer Expenditure estimates add behavioral context.
The internal AutoML pipeline package handles feature preparation, model selection, validation, evaluation, and experiment tracking so a method can be reused across audience definitions.
The classification path predicts likely audience membership from shared variables. The similarity path ranks lookalikes using distance and overlap when a clean target class is not available. Both workflows use a proprietary scoring and recommendation function: machine-learning evaluation metrics inform a custom score with penalties, helping select models and rank lookalikes for cross-survey use.
The internal audience-creation package carries the chosen method into a structured, validated, platform-ready profile, including the data transformation and ETL workflow needed to create audiences in the platform. Recurring production and deployment keeps those audience definitions current, while sizing combines survey incidence with ACS population data. The customized Consumer Expenditure ETL transforms BLS datasets into readable tabular expenditure data, so audience predictions can be applied on top.
The system turned questions about audience traits, media behaviors, and locations into reusable client audience definitions. More than 5,000 audiences were modeled for Fortune 500 clients, contributing 23% of Upsell MRR in H1 2025.