Top 3 · Customer support
- 01Gemini 3.8 FlashGoogle$1.50 USD / M tokens
- 02GPT-5.6 TerraOpenAI$4.50 USD / M tokens
- 03GPT-5.6 SolOpenAI$8.00 USD / M tokens
Control cost per conversation, product search and sales peaks
In e-commerce, AI is a volume game. A chatbot answering thousands of daily questions about shipping, sizing and returns turns every fraction of a cent per conversation into a visible budget line. Product search with embeddings gets better when shoppers describe what they want in their own words, but indexing and querying add cost too. And big sale days stress everything else at once, from the checkout to the model provider's rate limits.
There is no retail-specific AI benchmark, so we will not pretend there is one. What you can do is combine general LLM Stats indexes and official prices with your own conversations to estimate quality and cost. On the infrastructure side, the weight sits in image storage and its outbound traffic, in queues that never drop an order, and in a database that survives the busiest day of the year.
Pull a hundred real conversations and sort them by difficulty. Test a low-cost model and a stronger one, and measure how many each resolves correctly in every group plus the cost per full conversation, including history and catalog context. Many teams end up routing common questions to the cheap model and escalating the rest.
Embed titles, attributes and descriptions, then combine that with keyword search for exact SKUs, brands and sizes. Evaluate with real queries from your search logs and check whether the right product lands in the top ten. Re-embed only the products that change so you never pay for full reindexes.
In object storage with a CDN in front. Price per stored GB is rarely the issue; egress and per-request operations are. Before comparing providers, estimate GB served per month, how many variants you generate per image and what share the CDN answers from cache.
Write the order to the database and publish the event to a queue as one logical operation, for example with the outbox pattern. Inventory, payments and notifications consume at their own pace and retry with a per-order idempotency key. Watch the age of the oldest message throughout the campaign.
Load test at expected traffic plus a generous margin and see what breaks first: database connections, the model API's rate limits or checkout. Request quota increases weeks ahead, cache the catalog and decide which features you switch off if p95 latency crosses your threshold.
Tell us your volume and the options you are weighing. We reply in writing with the numbers of your real usage; no commitment.
No provider pays for its position. Indexes come from LLM Stats; prices from each provider's standard API. How we measure
Analysis, guides and new technology comparisons.