Pay Survey Data01

Pay survey data that cleans itself and sells year-round.

Shopping for benchmark data? This isn’t that — we build the operation behind the publishers who sell it: your survey, your methodology, encoded.

  • AI implementation services
  • Your survey methodology, encoded
  • Cleaning runs itself every cycle
  • A year-round product, not an annual PDF

Official services partner of the platforms defining AI

NVIDIAAnthropic
Peter Enestrom, founder of Zaigo

Every engagement is led personally by Peter Enestrom and the Zaigo AI & engineering team

YaleColumbia UniversityMicrosoft
The annual crunch01

The survey is the product. The pipeline is still Excel.

From the research directors and publisher-CEOs we have sat with: associations, membership orgs, and niche research publishers whose survey is the revenue product — and whose cleaning pipeline is one person’s spreadsheet.

  • Cleaned by hand before it can sell

    Outlier removal, smoothing, and triangulation applied manually in Excel — 1,600 responses a cycle before the percentile tables go out.

  • The instrument lives in a spreadsheet

    The intake survey is maintained as an Excel version of the questions, kept in sync with the published tables by memory.

  • A flat PDF can’t answer “where do I stand?”

    The annual survey report ships static, so every subscriber interprets and applies the compensation survey data themselves — or asks you to.

  • Benchmarks without cohorts mislead

    A $250M company can be 80th percentile overall but 45th in its size band — salary surveys by industry mean little without cohort rules.

  • The raw file can never ship

    Respondents share pay data because the raw spreadsheet never leaves — equity data especially — so the data stays locked in Excel.

What manual costs02

One report a year is the cap, not the plan.

What the manual way looks like at a publisher whose survey moves through Excel once a year. Your numbers will differ — the first cycle we map puts figures on yours before anything gets built.

~1,600Survey responses cleaned and normalized by hand per cycle — outlier removal, smoothing, stratification
Once a yearHow often a manual pipeline can publish — the survey’s value capped at a single annual PDF
45th, not 80thA $250M company’s size-band rank versus overall when the benchmark ignores cohort rules

From anonymized engagements — associations, membership orgs, and research publishers running ~1,500–1,600-response cycles

How it works03

We encode your methodology.

Not a survey platform, not another benchmark database subscription. The cleaning rules your research team applies by hand become the pipeline.

  1. 01

    Map one survey cycle

    Intake instrument, cleaning rules, published tables — we trace one cycle end-to-end, from raw response to the tables subscribers buy.

    First cycle
  2. 02

    Encode the methodology

    Outlier thresholds, smoothing, stratification, cohort bands by size and industry, respondent-privacy rules — documented against data you already collect, yours to keep.

    Per methodology
  3. 03

    Percentiles as the product

    Cleaning runs itself; odd responses land in a review queue. Subscribers enter title, size, and industry — they get their band; raw responses never surface.

    Every cycle
Not another tool

Survey platforms collect. Benchmarking services sell their own data. Neither is your pipeline.

The survey platform collects the responses but doesn’t clean them. The benchmark giants sell their own data, not yours. Excel does the cleaning — by hand, once a year. None of them is the pipeline.

Our AI does the reading — messy responses, public filings, free-text job titles — while your encoded methodology does the judging: what counts, what gets smoothed, which cohort band each company lands in, what never ships raw.

In production
Every cycleResponses cleaned by encoded rules — outlier thresholds, smoothing, stratification — not by hand once a year
Aggregate onlyWhat the self-serve tool can ever return — percentile bands, never a raw response
Year-roundHow long the percentile engine sells — not the week the annual PDF ships
We cleaned sixteen hundred responses by hand before every report. Now the rules do it, and the survey sells year-round instead of once.
Research director, membership and benchmarking organization
Peter Enestrom, founder of Zaigo
Who builds it

Led by Peter Enestrom.

Founder — leads AI & Engineering

Pete Enestrom

Every engagement is led personally by Pete, working with the Zaigo AI & engineering team from the two-week audit through the production handover. The person who scopes the work is the person who builds it.

Education
Yale & ColumbiaGraduate
Background
Microsoft & IntelFormer
Experience
Exited FounderVenture-Backed

Background

Questions04

Asked by research directors and publisher-CEOs.

The straight answers, before you book anything.

Pay survey data is compensation information — salaries, bonuses, equity — collected directly from employers through a survey run by an association, membership organization, or research publisher, then cleaned and published as benchmark tables. Because it comes from the publisher’s own respondent pool rather than a purchased database, respondent-privacy rules govern what can ship.

Industry associations, membership organizations, trade publishers, and niche research firms — the operations behind a compensation survey provider. Running one means maintaining the intake instrument, cleaning every cycle’s responses — outlier removal, smoothing, stratification — and publishing the benchmark tables, usually once a year.

No. Respondents share pay data because the raw file never leaves the publisher. What buyers can receive is aggregate percentile bands filtered by cohort — company size and industry — so no individual response can be reverse-engineered. Encoding that privacy rule into the pipeline is what makes a self-serve tool safe to sell.

Public filings — named-officer pay extracted from DEF 14A proxies at roughly 5,500-company scale — are normalized against the same rules as the survey responses: same job-title mapping, same cohort bands. Two messy sources become one benchmark dataset, and the survey stays the differentiated half of it.

Every cycle the survey collects. A manual pipeline caps refresh at once a year because cleaning is the bottleneck; with the methodology encoded, percentiles recalculate as responses arrive, so the published benchmark never goes stale between annual reports.

Start with one workflow.

Tell us where your team loses hours. We will come back with a straight answer on whether AI can help, what it would take, and what it would pay.