8. Synthetic customers, part 1: can a fake shopper judge your recommender?
Picture a shop that sells colored and flavored toilet paper. Mint. Lavender. Bacon. You build a new “customers also liked” box. Offline, it scores well. You ship it to half the traffic. Three weeks later, the A/B test (a live split test) shows no difference.
That is the normal story. This post is about a faster check: fake customers. Let them click on your box first. It is part 1 of 2. Part 1 is what the research says. Part 2 is what I built.
[ THE PROBLEM: TWO BAD OPTIONS ]
A recommender suggests products. To test a new one, you have two options. Both have flaws.
Offline metrics replay old logs. Two examples are nDCG (normalized discounted cumulative gain) and precision@k (how many of the top k picks are right). They are fast and cheap. But a model can win offline and lose online. This is the offline-online gap. A 2022 survey (Stavinova et al.) lists four causes:
- The site changed after the logs were made.
- The metrics ignore novelty and variety.
- Feedback loops are missing. What you show changes what people click next.
- The logged data is biased.
Online A/B tests are honest but slow. As post 3 showed, a small gain needs lots of traffic and time. Real customers may see a worse page. Privacy rules also limit the data you can keep.
A simulator sits in the middle. You learn fake users from real data. You let them react to a recommender. You count what they do. You get an answer in hours. No real customer is touched.
[ 5 PARTS OF A SIMULATOR ]
The survey read 18 simulators. They all share five parts, called C1 to C5.
All 18 have C1 (make fake data) and C3 (run a recommender on it). About half have C2, the what-if part. “What if prices go up?”
Now look at C4. Only 4 of 18 check the simulator itself. Does the fake world match the real one? Almost nobody asks. Remember this. It is the point of this series.
A simulator learns from logs, so it learns their biases too. The survey lists five. Positivity: people rate what they like. Popularity: popular items get more clicks. Conformity: people copy others. Exposure: people only see what was shown. Position: slot 1 gets clicks for being first. Most simulators only measure these. The survey names SOFA (Simulator for OFfline leArning and evaluation) as one that corrects bias. Position bias returns in part 2. It is the biggest one.
[ FOUR WAYS TO FAKE A USER ]
Each paper fakes a different layer. One fakes the data. Two fake the person. One fakes a chat.
1. Fake data, no fake people
Outbrain researchers (Malenšek et al., 2024) make fully artificial tables. You pick the rules: values per column, distribution shape, how columns interact, how much noise. You wrote the rules, so you know the truth. Then you test if an algorithm finds it.
They hid feature pairs that matter only together. The rules were logic gates (AND, OR, XOR) and sum of squares. The pairs sat among hundreds of useless features. DeepFM (deep factorization machine, a neural recommender) beat logistic regression in 10 of 11 setups. In config 6, both were near chance, and the simple model won. Real data cannot teach you this, because nobody gives you the true answer.
The catch: this makes data, not shoppers. The paper also never compares its data to real data.
2. A language model plays one shopper
SimUSER (Bougie and Watanabe, Woven by Toyota) gives each fake user an LLM (large language model) brain. It has two steps.
Step 1: pick a persona. Take about 50 items a real user liked or disliked. The LLM drafts five personas (age, job, Big Five personality traits). Keep the persona that best separates this user from a random other user.
Step 2: run the agent. It has four modules. Persona sets pickiness and habits. Perception reads item thumbnails. Memory holds past ratings and a knowledge graph. The brain reasons, then picks CLICK, NEXT, or EXIT.
The check is the best part. The authors had 55 real A/B tests, scored by pages visited per session. Did the simulator rank the variants like reality? Its Spearman correlation (how well two rankings agree) was the highest. It beat Agent4Rec and RecAgent. Then they tuned a recommender two ways. Tuning for nDCG@10 (nDCG on the top 10) gave results on par with the baseline. Tuning on SimUSER improved engagement in production. The catch: the data is private.
3. Train a small model on real feedback
A click shows what happened, not why. UserMirrorer (Wei et al., 2025) adds the why. Each example is a scene: the user’s profile and history, plus the list of items shown. The model picks one item.
A big LLM writes a reason for each real choice. A small model learns from the choice and the reason. Two tricks keep it cheap. First, keep scenes where a weak and a strong model disagree most. Those teach the most. Second, check each reason against the real choice. Matching reasons become training data for SFT (supervised fine-tuning). Non-matching ones become negative examples for DPO (direct preference optimization). The result is a 3-billion-parameter model. It was tested on eight public datasets. The code and models are open.
4. Chatty fake users
UserSimCRS v2 (Bernard and Balog) covers conversational recommender systems (CRS). It supports two designs. The agenda-based user has clear steps: NLU (natural language understanding), a policy, and NLG (natural language generation). It also has a fixed information need. v2 adds LLMs for reading and writing text. The end-to-end user is one LLM. v2 also adds an LLM judge to score chats. The authors warn that a judge can be biased or gamed.
[ MICRO, MESO, MACRO ]
The survey names three scales for fake users:
- Micro: learn from one user’s history. Most users do not have enough history.
- Meso: learn from a group of similar users. One model per group.
- Macro: learn from everyone. One model. Most simulators do this.
Meso is rare. I picked it for part 2. Shop customers differ a lot. A hotel buys for 40 rooms. A gift buyer buys one roll. One average shopper fits nobody. But one model per person has too little data.
[ WHICH ONE SHOULD YOU USE? ]
It depends on what you have and what you want to test.
- You want to test an algorithm against a known truth. Use fake data, like the Outbrain tool.
- You want rich behavior, like reactions to pictures and reviews. Use an LLM agent, like SimUSER. It costs tokens on every run.
- You have lots of real feedback and want cheap runs. Train a small model, like UserMirrorer.
- Your recommender is a chatbot. Use UserSimCRS.
Whatever you pick, ask the C4 question first. How will I check that this fake user acts like a real one? If the answer is “I won’t”, you have a toy.
[ WHAT NOBODY HAS PROVEN ]
The survey ends with four open problems:
- Recommender rankings on fake and real data sometimes disagree.
- There is no standard way to validate a simulator.
- There is no head-to-head benchmark.
- Few papers show real business gains.
SimUSER’s authors add an ethics warning. Fake users can amplify bias. They can also be used to manipulate people.
So the real question is not “can an LLM act like a shopper?” The question is this: do fake customers rank three recommenders in the same order as real customers? That is the only test that matters. It needs real data.
Part 2 builds this without LLMs, on an invented toilet paper store. It has 650 fake shoppers, one per group of real customers. The numbers are made up. The mistakes are real, and I list every one.
[ REFERENCES ]
All on arXiv. Figures 3 and 5 are reproduced under their licenses (CC BY 4.0 and CC BY-SA 4.0, both Creative Commons). The rest are my redrawings.
- Stavinova et al. Synthetic Data-Based Simulators for Recommender Systems: A Survey. 2022. arXiv:2206.11338
- Malenšek, Škrlj, Mramor, Demšar. Generating Diverse Synthetic Datasets for Evaluation of Real-life Recommender Systems. 2024. arXiv:2412.06809
- Bougie, Watanabe. SimUSER. 2025. arXiv:2504.12722
- Wei et al. Mirroring Users. 2025. arXiv:2508.18142
- Bernard, Balog. UserSimCRS v2. 2025. arXiv:2512.04588
[ ELSEWHERE ]
GitHub · LinkedIn · Dev.to · Substack · 4thwithme.dev/blog
May the --force be with you. See you next week.