© 2026 Jeremiah Yee

How many Singaporeans is an AI worth?

7 Oct 2026

How well can AI predict what Singaporeans think? I tested 16 AI setups against real poll results to see how their guesses compared with asking a few actual people.

A few years ago, I built a social polling app called SenseUs. People answered questions, then got to see what everyone else thought. Should every company move to a four-day work week? Should Singapore allow vaping? How many kids would you like to have? Can money buy happiness?

About a thousand people answered, mostly between late 2022 and 2023. Then life got in the way of maintaining it, and the app went quiet.

But the answers were still there.

I've gone back to them before. In an earlier post, I looked at the happiness questions to see whether people really do settle back into feeling kinda average. This time, I had a different question: could an AI have guessed what everyone said?

Apparently this is now a business. Companies sell synthetic survey panels, and people are now making AI “twins” and “simulated populations” with checks against real survey answers. You ask an AI what people think before, or instead of, asking the people themselves.

I still had all those poll responses sitting around. Seemed like a good place to test this myself.

How do you measure a guess?

I wanted an answer that meant something. Saying a model has an error of X percentage points is fine, but what can you actually do with that?

So I measured it in “people” (or N).

Imagine asking 5 random SenseUs users a question and using their answers to guess how the whole crowd voted. On average, you'd be off by about 28 percentage points. Ask 20 people and that drops to about 16 points. Ask 40 and it's about 13.

Now give an AI the same question. If its guess gets as close as those 20 people, we can say it's “worth 20 people” (20N). At least for this particular job.

Each model only got the question, the answer options, and a short description of the audience, like “respondents to a Singapore community question platform”. Then it had to guess the split: “Yes 62%, No 38%”. It didn't get to see the real answers.

I used two sets of questions:

  • 399 SenseUs questions, each answered by at least 60 people. So 44,741 answers in total.
  • 132 public YouGov Singapore poll questions, from a panel designed to represent Singapore adults. A check that this wasn't just something peculiar about SenseUs's users.

I tested 16 AI setups across OpenAI, Anthropic, Google, xAI, DeepSeek and Alibaba. Many models let you choose how hard they think, so I picked settings of similar general capability. For example, GPT-6 Sol on its highest thinking setting scores about the same on a general intelligence index as GPT-6 Astra on its lowest.

About 20-odd people

Accuracy of 16 AI setups in equivalent real respondents: GPT-6 Astra leads at about 18 SenseUs users and 27 YouGov respondents.

The best model, GPT-6 Astra on low thinking, was worth about 18 SenseUs users and 27 YouGov respondents. GPT-6.1 Sol, Claude Fable 5.1, Gemini 3.8 Flash and Claude Opus 5.5 were close behind. The weakest setups were worth about 7.

That's a fairly useful guess from just reading a question. It's also a long way from a survey of 1,000 people. If someone offered me a panel of AI respondents, I'd want to know how many actual people it was worth.

The individual questions are more fun to look at.

Every company should implement 4-day work weeks.
Strongly disagree 1 · 2 · 3 · 4 · 5 Strongly agree

85% of SenseUs users agreed, and 68% strongly agreed. GPT-6 Astra and Claude Opus both put “strongly agree” at around 40%, spreading the rest across milder answers. The users were pretty sure but the models expected them to be more on the fence.

How many kids would you like to have?
None · 1 · 2 · 3 · 4 · 5 or more

67% wanted exactly two. Both models guessed about 40% and overestimated how many wanted none. Maybe they'd read too many headlines about Singapore's birth rate.

Singapore should allow vaping.
Yes · No

35% said Singapore should allow it and 65% said no. Claude Opus guessed exactly 35/65.

Have you used ChatGPT for your assignments or work yet?
Yes · No

Asked whether they'd used ChatGPT for assignments or work, 61% of users said yes. Both models guessed 60–65%. Note that this was in 2023, a dinosaur age ago.

Of course, it's easy to pick out the spectacular successes and failures. Only about one in ten of Astra's guesses came within 5 points of the real split. About one in ten missed by 25 points or more. Most were somewhere in between.

I also tried a baseline that ignored the question entirely and just predicted the usual split for that type of question. That was already worth about 6 people. Every model beat it, which is reassuring. Reading the question does help.

Surely a smarter model would do better?

General intelligence benchmark scores versus polling accuracy on SenseUs and YouGov; higher intelligence does not consistently predict better guesses.

Qwen3.8 Max scores fairly high on general intelligence benchmarks, but gave some of the weakest guesses here. Claude Sonnet 5.5, which was just released, scores higher on those benchmarks than the Sonnet before it, but it wasn't any better at guessing.

Asking models to think harder usually made little difference too, even though it cost more. Being good at reasoning doesn't seem to guarantee a good read on what people think.

What if we combined the brains of different models?

I tried averaging their predictions. It didn't beat the best single model. The problem is that they tend to get the same questions wrong in the same direction. On the four-day work week, Astra and Opus agreed with each other, and both underestimated “strongly agree” by 25 points or more.

Two AIs agreeing can feel reassuring. Here, they were just making the same mistake. Something to be aware of when using different models in everyday life.

Everyone's a bit too middle-of-the-road

Answer shares on SenseUs 1-to-5 scales: real users picked the middle 16.5% of the time, compared with 22.3% for the average AI prediction.

This was the clearest example. On SenseUs's 1-to-5 scales, real users picked the middle option, “3”, only 16.5% of the time. Every model expected more middle answers, around 20–25%.

On YouGov, most models also underestimated how often people said “Don't know”. OpenAI's models underestimated it the most. So the AI's idea of an average respondent was a little too ready to have an answer, and on SenseUs, a little too ready to sit in the middle.

I tried correcting for that, moving some of each SenseUs scale prediction away from the middle and towards the top, where users actually answered. I worked out the adjustment on one half of the questions and tested it on the other. It improved all 16 models, lifting the best from 18 to 21 people's worth of accuracy.

Promising, until I tried it elsewhere. On a separate set of SenseUs questions, it helped GPT-6.1 Sol much less. On YouGov, where the models already got the middle about right, it didn't help.

I'm still not sure why SenseUs was different. Users could skip any question, so perhaps the people who answered were mostly those with an opinion. YouGov's panel goes through every question, including the ones people don't feel strongly about.

Or maybe it's the format. SenseUs showed bare numbers from 1 to 5; YouGov writes its answers out in words. On the few YouGov scales with worded labels, people actually picked the middle more often than the models expected.

Sadly, my data can't tell those explanations apart.

Was the Singapore part difficult?

Questions about BTO flats, NS, 377A and CNY red packets weren't harder for the models than generic ones. On YouGov, they did especially well on questions rating the government's performance. Apart from those, the local questions looked much like everything else.

The bigger difference was how you asked. Models did best on rating scales and worst on yes/no questions. For yes/no questions, the typical model was worth about 8 SenseUs users. You'd have done better asking 10 real people.

GPT-6 Astra and Gemini 3.8 Flash held up across all the formats, so the choice of model matters most if you're asking yes/no questions.

There was one useful clue about when to be cautious: when the models disagreed with each other, they were somewhat more likely to be wrong. It wasn't a reliable way to decide which answers to trust, but it could help you decide where to collect more real ones.

What if you can ask a few people?

Combining real answers with GPT-6 Astra: 5 real answers plus AI are as accurate as about 23 people; 20 real answers plus AI are worth about 37.

This is where the AI became much more useful.

Say you're doing a student project, asking around in a community group, or running a small team poll. You can get a few answers, but nowhere near a thousand. You could start with the AI's guess, then let each real answer nudge it.

Using GPT-6 Astra's predictions, averaged across the SenseUs questions:

  • 5 real answers plus the AI were about as accurate as polling 23 people.
  • 20 real answers plus the AI were about as accurate as polling 37 people.

That's about 18 extra people's worth of accuracy in the best setup I tested. Every model helped at every sample size, and the best ones added at least as much on the YouGov polls. It isn't a promise about any AI you pick, but it's a useful starting point if you can only ask a handful of people.

I also tried showing the AI the first few answers and asking it to guess again. That did worse than the simple nudging. It overreacted to the handful of answers it saw.

Giving it results from other questions was a little more helpful. Adding three past SenseUs results to each prompt improved its guesses slightly, mostly on 1-to-5 scales. Random past questions worked as well as similar ones. On YouGov, it made no real difference.

But what if it had just seen the answers before?

For most real world AI tests, this question is the elephant in the room. I couldn’t rule it out so I checked for signs of memorisation.

SenseUs results were mostly behind a login. When I asked two of the top models about the 399 questions, they recalled none of the results. They hadn't even heard of SenseUs. At least we weren’t that popular.

For YouGov, whose results are public, two models recalled none of the 138 results they were asked about. The models also ranked in almost the same order on the 2025–26 polls as on older ones. Reassuring, but asking a model whether it remembers something isn't conclusive.

So I later tested 91 polls from a public Singapore Telegram poll channel, posted between July and September 2026. These came after the training cutoff of every model I tested on them, and I recorded the AI guesses before collecting the results.

Astra was worth 13 voters there, compared with 3 for the baseline that ignored the question. Lower than on SenseUs, probably partly because most of these polls had only two options, where the models struggled most. But the models ranked in almost the same order as on SenseUs, and their accuracy didn't drop after their training cutoffs.

One other result puts that 13 in perspective. When the channel reposted an old question, its previous result was worth 87 voters as a guess of the new one. So there’s still plenty of room to do better.

Coming back to the people

I started with a fairly odd question about an app I wasn't maintaining anymore. The answer turned out to be quite practical: a good model can give you a small poll's worth of accuracy from the question alone, and combining that guess with a few real answers gets you further.

What worked less well was throwing more AI at it. More thinking usually didn't help much. Averaging models didn't beat the best one. The correction that worked on SenseUs didn't carry over to YouGov, or even fully to other SenseUs questions.

In my old SenseUs posts, I wrote about wanting to hear different perspectives and get people talking. There's something a bit funny about returning to the same app to see how far I could get without asking anyone.

Apparently, about 20 people. I'd still want to hear from the people themselves though.

About the data

I built SenseUs, and its users chose to be there, so they don't represent all Singaporeans. The method and limitations are in a simple paper I'm writing.

89 of the 399 SenseUs questions were items from a series that didn't make sense on their own. Without them, the best model was worth about 22 users instead of 18, so the numbers here are on the cautious side.

No personal information was given to the AI models. They saw only the question, the answer options and a one-line description of the audience. The analysis used anonymised answer records to work out vote splits. Everything reported here is at the level of whole questions, not individuals.

If you're building or buying AI respondents, or you'd like to test a model against the SenseUs responses, get in touch.

Try it with your own question

I've also been thinking about how this idea could be useful, so I built it into First Guess: the AI makes the first guess, and real answers take over as they come in. If there's a question you've been meaning to ask, it's free to try.


Back to blog