·Validation·Minds Team

Can AI Recreate a Real Gen Z Food Survey? Minds Reached 93.99% Approximation

Across three meal-frequency questions and 21 answer-option cells, Minds averaged 6.01 percentage points of error, reached 78.98% distribution overlap, and completed every planned response.

Sign in to download PDF

A real UK survey asked young people how often they eat out or buy takeaway food for breakfast, lunch, and dinner. This outcome-blind applied validation asked the same questions to a locked audience of 301 Gen Z Minds and compared the aggregate response distributions with the published human results.

The primary result was 93.99% aggregate approximation to the real survey distributions. Across 21 answer-option cells, the average gap was 6.01 percentage points. All 903 planned Mind-question responses completed successfully.

The analysis evaluates aggregate distributional agreement. It does not treat the score as individual prediction accuracy or as evidence that the same performance will hold for every audience and question type.

The Result in One View

Validation resultValueWhat it means
Approximation to the real survey93.99%100% minus the average percentage-point gap across 21 answer-option cells
Average gap to the real survey6.01 ppThe average absolute difference between Minds and real survey percentages
Mean distribution overlap78.98%How much of each full answer distribution overlapped
Completed responses903 / 903All 301 Minds answered all three questions
Pearson correlation0.804Strong alignment across the 21 real and synthetic percentages
Spearman correlation0.834Strong agreement in how answer options ranked

Research Question and Reference Data

Meal frequency depends on routine, work or study, budget, convenience, social context, and local availability. It therefore provides a practical test of whether a synthetic audience can reproduce a full consumer-response distribution rather than only produce plausible qualitative comments.

The reference came from the UK Food Standards Agency's Food and You 2, Wave 10 study. Fieldwork ran from 9 October 2024 to 7 February 2025, and the public dataset was issued on 25 September 2025. The analysis used the published results for people aged 16 to 24, with a human question base of 257 for each evaluated item.

The three registered questions were:

  • How often do you eat out or buy food to take out for breakfast?
  • How often do you eat out or buy food to take out for lunch?
  • How often do you eat out or buy food to take out for dinner?

Each question had seven registered answer options. The confirmatory analysis therefore compared 21 synthetic and human percentages.

The Audience: 301 Persistent Gen Z Minds

The synthetic panel represented young people aged 16 to 24 across England, Wales, and Northern Ireland. The locked cohort contained variation in age, gender, nation, income, routines, and social context. All 301 Minds were selected before the current outcomes were inspected.

Audience composition

Age band
  • 1
    16–1829%
  • 2
    19–2137%
  • 3
    22–2434%
Gender
  • 1
    Female65%
  • 2
    Male35%
Nation
  • 1
    England84%
  • 2
    Wales8%
  • 3
    Northern Ireland8%
Food and You 2 Survey: Wave 10 dataset
Food and You 2: Wave 10 research report

These were persistent audience members rather than temporary characters generated for a single question. The panel was grounded in 20,500 knowledge items represented by 14,671 retrieval chunks, allowing the same synthetic cohort to be reused across research tasks.

Study Design and Estimands

The study used a preregistered, same-cohort replication design:

  1. Lock the 301-person synthetic audience.
  2. Ask the original survey questions with the original answer options.
  3. Keep the real survey percentages out of the Minds runtime.
  4. Aggregate all 903 Mind answers into three complete distributions.
  5. Compare every synthetic answer-option percentage with the published human result.

Rather than forcing uncertainty into a single hard vote, Minds represented how likely each audience member was to choose each available answer. These probabilities were then combined into the final panel distribution.

Design elementRegistered specification
Synthetic sample301 persistent Minds aged 16–24
Human referencePublished Food and You 2 youth distributions; base n=257 per question
Instrument3 meal-frequency questions × 7 options
Planned responses903 Mind-question responses
Primary endpointUnweighted mean absolute percentage-point error across 21 cells
Secondary endpointsDistribution overlap, Jensen–Shannon divergence, Pearson r, and Spearman rho
Outcome isolationHuman percentages and benchmark files withheld from the runtime

Percentage approximation is defined as 100% − mean absolute percentage-point error. Mean distribution overlap is 1 − total-variation distance, averaged across the three questions. Jensen–Shannon divergence measures distribution-shape difference in bits, where zero indicates identical distributions.

Every Question Scored Above 91

Average gap to reality

Reported results for this study. The table and discussion retain the study scope, baselines and uncertainty.

  • Breakfast frequency2.98
  • Lunch frequency6.48
  • Dinner frequency8.56
0Percentage points · lower is better10
QuestionApproximation to the real surveyAverage gap to reality
Breakfast frequency97.02%2.98 pp
Lunch frequency93.52%6.48 pp
Dinner frequency91.44%8.56 pp
All 21 answer-option cells93.99%6.01 pp

Breakfast had the lowest cell-level error, while dinner had the highest. The aggregate result was not driven by one question alone: each question achieved more than 91% approximation to its real survey distribution.

Secondary Distributional Checks

Because mean absolute error alone does not fully describe distribution shape, the analysis included four complementary checks.

  • 78.98% mean distribution overlap: most of the response mass appeared in the same places in the real and synthetic results.
  • 0.0602 Jensen–Shannon divergence: the overall shapes were close; zero would mean identical distributions.
  • 0.804 Pearson correlation: large and small percentages tended to move together.
  • 0.834 Spearman correlation: answer options were ranked similarly by prevalence.

All four metrics point in the same direction: within this instrument, the Minds distributions were close to the human reference in both absolute level and relative shape. Correlation is reported as a companion measure and does not replace level-error metrics.

Response Completeness and Grounding

Every one of the 903 accepted answers carried provenance:

  • 889 answers used retrieved Mind knowledge;
  • 14 answers used the Mind's intrinsic profile;
  • 0 answers were marked unsupported; and
  • 301 of 301 Minds completed every question.

These provenance fields establish that the run used the persistent Mind pathway rather than unsupported answers. They do not by themselves prove that every retrieved memory was causally necessary for the final distribution.

Practical Interpretation

Within the tested scope, the result supports using a grounded synthetic panel for rapid, iterative research such as:

  • test campaign claims before media spend;
  • compare product or packaging concepts;
  • explore why different audience segments react differently;
  • refine survey questions before fieldwork; and
  • identify the strongest directions for later human validation.

The result supports fast directional and comparative work. High-stakes claims, regulated decisions, precise market sizing, and questions outside the validated scope may still require recruited respondents and an appropriate human study design.

Scope and Limitations

This was an independent Minds validation using public Food and You 2 data. The Food Standards Agency produced the reference survey but did not sponsor, endorse, or review the Minds study.

This study supportsThis study does not establish
Close aggregate agreement on three UK youth meal-frequency distributions93.99% individual-level prediction accuracy
Complete production execution for 301 Minds and 903 responsesUniversal performance across audiences, countries, domains, or question types
Agreement across MAE, overlap, divergence, and correlation metricsCausal proof that grounding alone produced the observed agreement
Outcome isolation at runtimeComplete exclusion of indirect exposure to a public survey through pretraining or the persistent knowledge corpus

This is a same-cohort, same-instrument replication covering three closely related questions from one survey. The human and synthetic respondents were not individually matched, and the analysis compares aggregate distributions. No sampling-uncertainty interval from the human microdata was estimated for this public result. Further validation on disjoint audiences and instruments is required before making a broader accuracy claim.

Sources

Frequently asked questions

How closely did Minds match the real Food and You survey?

Across 21 answer-option percentages, Minds achieved 93.99% aggregate approximation to the real survey. The average gap was 6.01 percentage points, and mean distribution overlap was 78.98%.

Does 93.99% approximation mean Minds predicts every individual correctly?

No. The 93.99% approximation is calculated as 100% minus the average percentage-point gap across 21 aggregate survey cells. It measures how closely the audience distribution matched the real survey, not whether every individual answer was predicted correctly.

Were these real respondents?

No. They were 301 persistent synthetic audience members aged 16 to 24. Their aggregate answers were compared with the published results of a real UK Food Standards Agency survey.

Were the real survey results shown to the Minds?

No. The questions and answer options were reproduced, but the benchmark dataset and real percentages were held out from the runtime. Because the source is public, indirect exposure through model pretraining or the persistent knowledge corpus cannot be excluded completely.

What makes this different from prompting a general AI 301 times?

The Minds were persistent audience members grounded in 20,500 knowledge items and 14,671 retrieval chunks. In this run, 889 of 903 answers used retrieved Mind knowledge, 14 used the intrinsic profile, and none were unsupported.