Williams-Sonoma spent its Q2 fiscal 2026 earnings call talking, at length, about a chatbot. Revenue associated with the Olive assistant on the flagship brand grew 620 percent. Engagement grew 700 percent. Customers who used it converted at three times the rate of shoppers who did not. Its sister assistant on the Pottery Barn family, Otto, resolved more than 70 percent of the conversations it handled without escalation. Numbers like that are usually a vendor slide, not an earnings call, so it is worth reading them slowly.
Chief Technology and Digital Officer Sameer Hassan tied the results to the company's AI tooling across sales and support, and the disclosure was clean enough that finance media covered it as a real earnings item, not a press release. Digital Commerce 360 wrote it up the day of the print. That earned it a slot in this week's news cycle, which means it will now appear in twenty decks by Friday, most of which will pull the wrong lesson.
The three numbers behind the 620 percent claim
Start with what was said, exactly. On the fiscal Q2 call the company said revenue associated with Olive grew 620 percent year over year and that engagement grew 700 percent since the start of the fiscal year. It also said Olive users converted at three times the rate of other customers. Otto, the newer Pottery Barn assistant, resolved more than 70 percent of its conversations without transferring the customer to a human. Separately, the company said visits that used its personalized experiences generated roughly nine times the revenue of an average visit, compared with about two times last year.
Every one of those figures compares a self-selected subset to a base. That is not a criticism, it is how retail assistants get measured. But it is why the framing has to be exact.
| NUMBER | WHAT IT ACTUALLY COMPARES | WHAT IT DOES NOT PROVE |
|---|---|---|
| 620% revenue growth for Olive | This year's Olive-attributed revenue vs. last year's, when the assistant was newer and smaller | That Olive caused 620% more revenue overall, or that removing Olive would cost that much |
| 3x conversion vs. other customers | Shoppers who engaged with Olive vs. those who did not, on the same site | That Olive lifted conversion; users who ask a shopping assistant were already higher-intent than average |
| 9x revenue per personalized visit | Visits using personalized experiences vs. an average visit across all traffic | That personalization created the 9x; the personalized experience gates itself to higher-intent behavior |
| 70% of Otto conversations self-resolved | Otto sessions that ended without a human handoff | That customers were satisfied; a session ending is not the same as a question answered |
None of this makes the numbers fake. All four are internally consistent with a retailer that has built a scoped assistant, put it in front of the right shoppers, and given it a real product graph to answer from. What they are not is a return-on-investment figure. Nobody on the call said Olive added N basis points to conversion at constant traffic. That is a different sentence, and it is the sentence a lot of AI-assistant vendor pitches quietly borrow from Williams-Sonoma without earning.
The confusion matters because the mechanic that Olive represents will get sold to your buying committee this quarter by three different vendors, and each of them will show a version of the same slide. Being crisp about which figure is a real lift, and which is a selection effect, is what keeps a legitimate program from being oversold and then killed after two quarters. The gap between what a chart shows and what a chart proves is the same measurement problem we walked through in the error bars framework for AI search reporting.
Why the conversion lift is a product story, not a marketing story
Olive works because it sits inside a very specific set of conditions.
The catalog is finite and structured. Williams-Sonoma sells knives, sheets, cookware, and furniture. Every SKU has attributes, images, price points, and reviews. That gives the assistant a real product graph to reason over, and it means answers can be grounded in a schema rather than generated from prose. A retailer that took the Olive playbook and pointed it at a poorly-tagged CMS would ship the same interface and get a fraction of the outcome.
The assistant is scoped tightly to a job. Olive does not answer general questions about kitchen design theory. It helps a shopper narrow a purchase from a known set of options and finish it. The scope is what allows the model to be usefully wrong less often, and it is what keeps abandonment rates on the assistant lower than a general-purpose chat would produce.
The other reason the story worked was that Williams-Sonoma already had a very high-context site. Product pages have real photography, real dimensions, and cross-linked reviews. That means the assistant does not have to fabricate a catalog to sound helpful, and it means the failure modes are honest ones (the product does not exist, the size is out of stock) rather than embarrassing ones (invented specs). The same investment in structured product data is why the company already ranked well in search before any of this, and it is why the assistant benefits from what is essentially the same content model we described in how AI engines decide what to cite.
What the mechanic looks like translated to B2B
A B2B software vendor is not selling saute pans, but the shape of the play is portable if you strip the retail vocabulary off it.
The parts that transfer are the parts that were structural: a scoped assistant, sitting on top of a clean product graph, that answers a narrow set of buyer questions with the actual catalog behind it. That is not a chatbot on a marketing homepage. It is closer to an interactive comparison surface that lives on your pricing page, your integrations page, and your solutions pages, and answers the kinds of questions a real buyer types into a general engine anyway. The equivalent of the Olive catalog for B2B is your integrations directory, your compliance matrix, your industry pages, and the actual product tiering behind them. Most of that data already exists and is not being used in the interface.
The reason the scoped-on-a-conversion-page pattern works in B2B is not that buyers want to chat with your website. It is that a comparison buyer will type the same three questions into ChatGPT or Perplexity if you do not answer them on the surface. If those questions are answered on your page, in your voice, grounded in your actual product, you keep the session. If they are not, the engine writes an answer that averages your competitors and yours, and the buyer decides from that. This is the same reason we push for real feature and pricing pages instead of gated PDFs in the generative engine optimization work we run for B2B software.
Where the transfer breaks, and why
The Williams-Sonoma model breaks at three predictable places when it hits an enterprise buying committee, and pretending it does not is how a program gets built and killed.
| RETAIL ASSUMPTION | WHAT BREAKS IN B2B | WHAT TO BUILD INSTEAD |
|---|---|---|
| One shopper, one session, one purchase | Five stakeholders across six weeks, needing different answers at different depths | A durable, sharable answer surface (comparison pages, integration matrix, ROI worksheet) that the champion can send to the CFO, not a chat that resets |
| Assistant closes the purchase in-session | Purchase runs through security review, procurement, and legal | Assistant collects the questions and hands off a scoped, referenceable answer bundle to the account team, not a fake close |
| Price and availability answer most objections | Compliance, data residency, SLAs, and integration depth answer most objections | Grounding on the compliance and integration corpus, not the marketing corpus; a wrong SOC 2 answer is a lost deal |
| Handoff avoidance is a cost win | Handoff avoidance can be a pipeline loss if the deal was worth the human's time | Score conversations for account fit and hand off high-fit accounts to a human on purpose, not by exception |
The last row is the one B2B teams get wrong most often. In retail, deflection is a good thing because the incremental margin on a $180 sheet set does not justify a human interaction. In B2B, a $180,000 account you deflected to a self-serve answer is a lost enterprise deal, and the report will show it as a support cost saved. The instrumentation has to run the other direction: score the account behind the session, and if it looks like a real buyer, escalate to a human before the assistant finishes its answer. That is the opposite of the Otto 70 percent number, and it is the correct choice.
The compliance row is the one where an unscoped assistant becomes an actual liability. A hallucinated bedding dimension costs a return. A hallucinated SOC 2 status kills a deal and produces a paper trail you cannot walk back. If you do build one of these, the grounding source for compliance answers cannot be the same corpus as the grounding source for feature answers. Two indices, two prompts, two guardrails. This is the same reason enterprise GEO readiness treats regulated content as a separate content class rather than a subset of marketing.
Do this quarter
The mistake to avoid is reading the Williams-Sonoma number as an argument to buy a general-purpose chatbot and put it on your homepage. That is not what they built and it is not what worked. The version of this that is worth doing in the next 90 days looks smaller and more specific.
Pick one high-intent page. Pricing, integrations, or the industry page for your largest vertical. Instrument what visitors actually search for on it, or what they type into a general AI engine right after they leave it. Those two sources give you a real question set, not a hypothetical one. If you do not have that data yet, this is what our reporting and analytics work sets up first, before any assistant is even scoped.
Then build the smallest possible assistant that can answer those questions from your actual product data. Not a general chat, not a support tool, and not the kind of splash animation vendors demo. A narrow answer surface, grounded in a small, versioned corpus, that either resolves the question with a real answer or hands off to a human with the conversation attached.
Measure it honestly. Matched cohort against non-users, not site average. Track the assist rate on wins, not just the deflection rate on losses. Publish the delta internally on a monthly cadence, and be willing to turn it off if the honest number does not clear the cost of running it. Most programs that get killed this way get killed because they were sold on the vanity metric and never had a fallback report ready.
The Williams-Sonoma story is a good story. The 620 percent is a real number. But the reason it worked is that they did the boring work first: clean catalog, scoped job, matched measurement, honest reporting up the chain. That is the part that ports. The chat interface is the least interesting variable in the whole thing.
So read the headline. Then read the sentence under it. Then build the version of it that fits how your buyers actually decide, not the one that fits how a home goods shopper picks a duvet cover. If you get that split right you will have something to report next year. If you get it wrong you will have a demo that was in the deck once. The gap between those two outcomes is the same content and measurement discipline that shows up in every B2B engagement we run, and it is what makes the difference between an AI program that survives its second board review and one that does not.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.