


Boosting sounds like a pragmatic solution to a familiar problem: sample sizes that are too small in important segments.
Classic boosting means deliberately surveying a subgroup more heavily, for example people over 70 or customers in a particular region. Synthetic boosting replaces this additional surveying: a model learns from the existing real answers and generates additional answers for the under-represented group.
The method therefore assumes that the model knows the group well. That is precisely what is often difficult with small, specific groups. Wang and colleagues showed in 2025 that language models portray demographic groups in a distorted and overly flat way. Brand and colleagues found that models only deliver usable values for new product attributes once they have earlier human data from the same category.
Synthetic boosting fills small subgroups, but it creates no new evidence. The additional cases contain no information that was not already in the real answers and the model. They make an analysis possible, but they do not make it more reliable, even if the sample size looks larger.
For decisions that depend on a segment, such as whether a tariff for the self-employed is worthwhile, this is an important difference. A seemingly solid base of 300 cases, 250 of which are simulated, deserves the confidence of 50 cases, not of 300.
A bank surveys 1,200 people about a new current account. Among them are only 45 students, a target group the product is aimed at. With synthetic boosting, the group is filled up to 300. The report shows 52 % interest among students with narrow confidence intervals.
The narrow intervals reflect the simulated cases, not additional certainty. Anyone who has to decide for this target group gains more from a test that reaches students specifically via ads and measures their sign-up intent.
Weighting changes the influence of real cases but does not create new ones. Imputation fills individual missing values for real respondents. Synthetic boosting creates entire additional respondents. Silicon sampling replaces a survey completely, boosting supplements it.
Boosting can be useful for exploratory analyses, for example to form hypotheses about a subgroup. Always report synthetic cases separately and calculate uncertainty on the basis of the real cases.
A behavioural test also has a limit here: it needs a target group that can be reached via ads. Very small, hard-to-address segments need more time or budget. Horizon clarifies this in advance in the test design.
Questions to check in a report with boosting: how many real cases lie behind each segment? Was the uncertainty calculated on this basis? And were the synthetic cases generated from data of the same category and the same country?
Wang, Morgenstern & Dickerson 2025: In studies with 3,200 participants from 16 demographic groups and four language models, the models portray groups in a distorted way and represent their internal diversity too flatly. Large language models that replace human participants can harmfully misportray and flatten identity groups, Nature Machine Intelligence. Source
Brand, Israeli & Ngwe 2025: A language model reproduces willingness to pay from conjoint studies for existing attributes; for new attributes, it needs fine-tuning with earlier human data from the same category, while data from other categories hardly helps. Using LLMs for Market Research, MSI Working Paper. Source
No. Weighting works only with real cases, boosting adds simulated cases.
Uncertainty should be based on the real cases. Simulated cases narrow intervals without providing new information.
For exploration and hypotheses, when synthetic cases are clearly labelled and no decision rests on them alone.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more