Rank first, trust later: Using LLM agents for retail experiments
Synthetic customers ranking different promotion cues tl;dr (never ai;dr) LLMs can simulate shoppers and help data science teams prescreen options before running experiments. One good use case is ranking options, not estimating effect sizes, if the model's bias is order-preserving. Simulation quality depends on methodological details such as persona construction and representation, clear experimental instructions, response elicitation, and robustness checks. Recent evidence shows a major robustness problem: even when persona content stays the same, changing only the textual representation can shift simulated agent behavior by more than 70 percentage points. At this point, synthetic panels are safest as exploratory playgrounds unless results survive serious robustness audits across prompts, personas, and model choices. Podcast-style summary by NotebookLM We are taking a break in our augmented data science series to ...