Pay survey data that cleans itself and sells year-round.
Shopping for benchmark data? This isn’t that — we build the operation behind the publishers who sell it: your survey, your methodology, encoded.
Official services partner of the platforms defining AI
The survey is the product. The pipeline is still Excel.
From the research directors and publisher-CEOs we have sat with: associations, membership orgs, and niche research publishers whose survey is the revenue product — and whose cleaning pipeline is one person’s spreadsheet.
Cleaned by hand before it can sell
Outlier removal, smoothing, and triangulation applied manually in Excel — 1,600 responses a cycle before the percentile tables go out.
The instrument lives in a spreadsheet
The intake survey is maintained as an Excel version of the questions, kept in sync with the published tables by memory.
A flat PDF can’t answer “where do I stand?”
The annual survey report ships static, so every subscriber interprets and applies the compensation survey data themselves — or asks you to.
Benchmarks without cohorts mislead
A $250M company can be 80th percentile overall but 45th in its size band — salary surveys by industry mean little without cohort rules.
The raw file can never ship
Respondents share pay data because the raw spreadsheet never leaves — equity data especially — so the data stays locked in Excel.
One report a year is the cap, not the plan.
What the manual way looks like at a publisher whose survey moves through Excel once a year. Your numbers will differ — the first cycle we map puts figures on yours before anything gets built.
From anonymized engagements — associations, membership orgs, and research publishers running ~1,500–1,600-response cycles
We encode your methodology.
Not a survey platform, not another benchmark database subscription. The cleaning rules your research team applies by hand become the pipeline.
- 01
Map one survey cycle
Intake instrument, cleaning rules, published tables — we trace one cycle end-to-end, from raw response to the tables subscribers buy.
- 02
Encode the methodology
Outlier thresholds, smoothing, stratification, cohort bands by size and industry, respondent-privacy rules — documented against data you already collect, yours to keep.
- 03
Percentiles as the product
Cleaning runs itself; odd responses land in a review queue. Subscribers enter title, size, and industry — they get their band; raw responses never surface.
Survey platforms collect. Benchmarking services sell their own data. Neither is your pipeline.
The survey platform collects the responses but doesn’t clean them. The benchmark giants sell their own data, not yours. Excel does the cleaning — by hand, once a year. None of them is the pipeline.
Our AI does the reading — messy responses, public filings, free-text job titles — while your encoded methodology does the judging: what counts, what gets smoothed, which cohort band each company lands in, what never ships raw.
We cleaned sixteen hundred responses by hand before every report. Now the rules do it, and the survey sells year-round instead of once.
Asked by research directors and publisher-CEOs.
The straight answers, before you book anything.
Pay survey data is compensation information — salaries, bonuses, equity — collected directly from employers through a survey run by an association, membership organization, or research publisher, then cleaned and published as benchmark tables. Because it comes from the publisher’s own respondent pool rather than a purchased database, respondent-privacy rules govern what can ship.
Industry associations, membership organizations, trade publishers, and niche research firms — the operations behind a compensation survey provider. Running one means maintaining the intake instrument, cleaning every cycle’s responses — outlier removal, smoothing, stratification — and publishing the benchmark tables, usually once a year.
No. Respondents share pay data because the raw file never leaves the publisher. What buyers can receive is aggregate percentile bands filtered by cohort — company size and industry — so no individual response can be reverse-engineered. Encoding that privacy rule into the pipeline is what makes a self-serve tool safe to sell.
Public filings — named-officer pay extracted from DEF 14A proxies at roughly 5,500-company scale — are normalized against the same rules as the survey responses: same job-title mapping, same cohort bands. Two messy sources become one benchmark dataset, and the survey stays the differentiated half of it.
Every cycle the survey collects. A manual pipeline caps refresh at once a year because cleaning is the bottleneck; with the methodology encoded, percentiles recalculate as responses arrive, so the published benchmark never goes stale between annual reports.
Start with one workflow.
Tell us where your team loses hours. We will come back with a straight answer on whether AI can help, what it would take, and what it would pay.


