Your redesign should fail 10,000 times before a real customer ever sees it.
Shopify showed SimGym as a research preview in the Winter '26 edition. It simulates shopper behavior against your actual store. Not a survey panel, and not a session replay of what already happened. Synthetic shoppers with different budgets, devices, and patience levels try to complete a purchase, then report where they gave up.
The work is in cohort design. A visitor on a 3 year old Android with a $40 ceiling. A returning buyer hunting 1 specific size. Someone who abandons after 2 confusing steps. Point them at the staging build and read the failure log. Then validate the fixes against real traffic with Rollouts A/B testing. Simulation generates candidates. Live tests decide.
Then publish the results. A post naming the exact step where 10,000 simulated shoppers quit your checkout, and what you changed in response, is more useful than anything else you could write that quarter. The transparency is the content, and almost nobody in your category will do it.
The honest catch: synthetic shoppers do not have your customers' context. They will not reproduce brand loyalty, gift panic, or a size chart your buyers have learned to distrust. A preview tool models plausible behavior, not your demand. Treat the output as a bug list rather than a forecast, and treat any predicted lift as a hypothesis until real traffic votes on it.
QA has always answered whether the store works. This answers whether the store is worth the effort of using. Those are different questions, and only 1 of them shows up in revenue.
