Product Feed Optimization for AI Shopping: How to Structure Your E-commerce Catalog for ChatGPT, Google, and AI Search
Product feed optimization used to be a channel problem. AI shopping is turning it into an architecture problem.
Customers can now ask ChatGPT, Google AI Mode and other AI systems to find, compare and narrow products before they ever reach a storefront. That means the quality of the answer increasingly depends on whether those systems can retrieve accurate product and commercial data from your business.
This article looks at what OpenAI and Google are actually asking merchants to structure, where retrieval breaks in the real world, and how to test your own catalogue for blind spots.
The core question: can an AI system retrieve an accurate, structured and current representation of what you sell?
We are deliberately focused on retrieval here. AI-side browsing and agentic commerce are the next layer, and we are working on a separate article for that.
What we mean by AI retrieval
Before going deeper, it helps to separate two things that are often grouped together under “AI commerce.”
- Retrieval: an AI system finding, understanding and answering questions about products, pricing, availability, compatibility and fulfilment.
- AI-side browsing and agentic commerce: an AI opening a website, navigating interfaces and, in some cases, taking actions on a customer's behalf.
This article is about the first.
A shopper can now ask:
“I’m in Richmond, Victoria and need a pool robot under $1,000 by tomorrow. Where can I get one, pickup or delivery?”
That is not just a product search. It combines a category, budget, location, deadline, fulfilment preference and an implied requirement for current inventory.
To answer it well, an AI system needs to retrieve more than a product title and description. It may need to resolve the right variant, current price, local availability, pickup eligibility and fulfilment timing.
That is why this article goes beyond conventional feed optimization. We are looking at whether the commercial information a business already has is actually available through a retrieval path an AI system can use.
A browser-capable agent may be able to interact with a website differently, including executing JavaScript or clicking through a store-availability experience. That is a separate technical problem, and we will cover it in the next article rather than mixing the two here.
There is a lot of commentary about “optimizing for AI shopping,” but very little of it separates documented platform requirements from speculation about ranking.
For this research we focused on information merchants can actually control.
Research scope
Platform specifications · Merchant guidance · Live retrieval testing
1. Platform specifications. We reviewed OpenAI's current stable product-feed documentation for ChatGPT discovery, including its supported product identity, variant, attribute, price, availability, shipping, returns and review fields. We also reviewed Google Merchant Center's product data specification, conversational attributes, local inventory data and current AI performance reporting.
2. Adjacent platform guidance. We reviewed Perplexity's current guidance on how product detail depth influences its recommendation system.
3. Live retrieval testing. On 2 September 2026 we ran a logged-out ChatGPT shopping test using a real Core dna customer, Clark Rubber. We deliberately used a query that required product discovery and local fulfilment information.
We then inspected the Clark Rubber product experience itself to compare what the business knew, what a human shopper could retrieve and what ChatGPT could retrieve in that test.
Important limitation: OpenAI, Google and other AI-shopping platforms do not publish complete ranking algorithms. Nothing in this article should be read as “add this field and you will rank.” Product recommendations can draw on query context, public web content, merchant feeds, third-party providers, reviews, price, availability and other signals.
Our research question is narrower: what commercial information are AI-shopping systems asking merchants to structure, and how reliably can those systems get to the information that already exists?
The easiest way to see the change is to look at what the current platform specifications actually ask merchants to provide.
OpenAI's stable product discovery feed starts with nine required fields for each purchasable product or variant: a stable item ID, title, description, product URL, brand, seller name, image URL, current price and current availability.
It then supports additional data for identity and variants, including group_id, variant_dict, offer IDs, GTIN and MPN, as well as product attributes such as material, colour, size, dimensions and weight. The stable discovery contract also includes shipping, returns and product review aggregates.
There is an important boundary in the current documentation: raw product Q&A lists are not part of OpenAI's stable discovery contract, and the stable schema does not currently expose the Google-style related-product relationship model or store-pickup fields we discuss elsewhere in this article.
Google's current Merchant Center model goes further in some of those areas.
Its optional conversational attributes include:
question_and_answerdocument_linkrelated_productitem_group_titlevariant_optionpopularity_rank
Google describes these attributes as a way to help AI systems and conversational agents understand product nuances. Its related-product field can explicitly represent relationships such as a required part, accessory, substitute or something often bought with the product. Its document-link field can connect authoritative PDFs such as manuals and assembly instructions to the product.
This difference between OpenAI and Google is not a problem to solve by choosing one schema.
It is actually one of the strongest arguments for having a better underlying product model.
Your internal commercial model should be richer than any one channel's current feed specification.
| Commercial information | OpenAI stable discovery feed | Google Merchant Center |
|---|---|---|
| Stable product / variant identity | Yes | Yes |
| Current price and availability | Required | Yes |
| Variant grouping | group_id + variant_dict | item_group_id + variant attributes |
| Colour / size / material | Supported | Supported |
| Shipping | Supported | Supported |
| Returns | Supported | Supported |
| Product review aggregates | Supported | Supported |
| Raw product Q&A | Not in the stable discovery contract | Optional conversational attribute |
| Authoritative product documents | No equivalent stable field | Optional document_link |
| Explicit related-product relationships | No equivalent stable relationship field | Optional related_product |
| Store-level inventory / pickup | Not part of the standard stable discovery upload | Available through Google's local inventory data model |
The exact contracts will continue to evolve. The architectural lesson is not to chase parity between them. It is to make sure your business can map the right commercial truth into whichever fields each channel supports.
Google's current AI performance reporting makes this shift unusually visible.
The live Merchant Center report covers conversational shopping queries on AI Mode and AI Overviews. It shows merchants metrics such as share of voice, products showing, top terms, popular attributes and top search intents. Google also uses an attribute-completeness view to highlight structured specifications shoppers are looking for, with examples such as size, colour and material.
That matters because conversational queries naturally combine more constraints than traditional search terms.
Compare “running shoes” with:
“I slightly overpronate, run mostly on pavement, want good cushioning without anything too soft, need a wide fit and want to stay under $160.”
The second query contains product intent and decision criteria. Add “women's 8.5, black, in stock and delivered by Saturday,” and the system also needs to resolve a purchasable variant, current inventory, price and fulfilment.

This is why a useful catalogue audit begins with a different question from a conventional feed audit.
Do not start with “Which fields does ChatGPT require?” Start with:
What does our business actually know about this product, and which of those facts could matter to a buying decision?
For a jacket that might be waterproof rating, breathability, insulation, fit and weight. For an industrial pump it might be voltage, phase, flow rate, certification, compatible fittings and operating temperature. For furniture it could be exact dimensions, material, finish and assembly requirements. For replacement parts, compatibility may be more important than every marketing attribute combined.
The goal is not to create hundreds of fields because AI exists. It is to identify the facts that explain what the product is, who or what it is for, how it differs, what it works with, what can be purchased now, and under what commercial conditions it can be fulfilled.
A useful catalogue audit lens
Identity · Attributes · Variant · Price · Inventory · Fulfilment · Compatibility
Perplexity's current merchant guidance reinforces the point. It says merchants that provide deeper product information such as availability, reviews, pricing and specifications are more likely to be recommended by its answer engine.
Structured data still does not guarantee recommendation. But thin product data gives a conversational system less explicit commercial information to reason against.
A lot of ecommerce information still exists primarily for presentation.
Colour may exist in a title. Dimensions may sit in a paragraph. Compatibility may be buried in a PDF. A specification table may be rendered by the CMS but not represented as typed product data. The storefront can still look complete because a human can read all of those elements together.
That is not the same thing as the commerce platform actually knowing the facts.
Consider this copy:
“Our most versatile shell yet, built for adventure whatever the weather.”
That may be perfectly good merchandising.
But if an AI assistant is trying to determine whether the product fits a specific use case, this is more useful:
waterproof_rating = 20,000 mm breathability = 15,000 g/m² material = recycled nylon insulation = none fit = regular weight = 410 g seams = fully sealed
Story + structure
The first tells the story. The second describes the product. You need both.
This distinction becomes even more important in categories where product fit depends on technical detail. If a pool cleaner supports floor and wall cleaning but not waterline cleaning, that is not a copywriting nuance. It is a purchasing constraint. If a spare filter fits the 2025 model but not the 2024 model, that is not a recommendation block. It is a compatibility relationship.
A machine can sometimes infer those things from prose. But inference is less reliable than explicit commercial data, particularly when the user adds several constraints at once.
Finding 4: Product identity and variants matter before optimization starts
Product-feed advice often jumps quickly to titles and descriptions. We would start earlier: identity.
Every product and purchasable variant needs a durable way to be reconciled across systems. That usually means a stable internal ID and SKU, with recognized identifiers such as GTIN, UPC, EAN or MPN where applicable.
Why does that matter in an AI-shopping environment? Because the same commercial object may be represented simultaneously in your storefront, Google Merchant Center, an AI feed, a marketplace, a mobile app, a store-inventory service and an internal sales tool. If those representations cannot be reconciled reliably, every downstream experience becomes harder to keep current.
Variants deserve particular attention. Customers rarely purchase an abstract parent product. They purchase something like Alpine Shell / Navy / Medium.
That variant can have its own SKU and GTIN, image, price, inventory, availability, pickup eligibility and delivery promise.
Parent product ≠ purchasable variant
Alpine Shell → Navy → Medium → SKU → Price → Stock → Pickup / delivery
OpenAI supports product/variant grouping and maps fields such as colour, size and material. Google recommends item-group identifiers and ProductGroup/hasVariant/variesBy relationships so that it can understand how purchasable options relate to the parent product.
That gives ecommerce teams a deceptively simple audit question:
Do we know that the product is available, or do we know that this exact purchasable variant is available?
The distinction matters as soon as a user says “navy, medium, under $250, available before Thursday.” If the parent product says “in stock” but the requested variant is not, the catalogue is technically populated but commercially misleading.
Attributes tell a machine what a product is. Relationships tell it how products fit together.
This matters because conversational commerce naturally produces relational questions:
- Which filter fits this machine?
- What replaced the discontinued model?
- What else do I need to install this?
- Which accessory works with the product I already own?
- What is a cheaper substitute?
Google's current Merchant Center related-product specification gives merchants explicit ways to model this information for conversational experiences.
Its optional related_product attribute supports relationship types including:
part_of_setrequired_partoften_bought_withsubstitutedifferent_brandaccessory
Google also provides optional question_and_answer and document_link attributes. The latter can connect authoritative PDFs such as manuals, user guides and assembly instructions to the product so conversational systems can use them during research.
OpenAI's current stable discovery feed does not expose the same relationship and raw Q&A contract. Its documentation explicitly says raw question-and-answer lists are not part of the stable discovery contract.
That difference is useful.
It shows why your underlying catalogue should not simply mirror one destination schema. The business may know that two products are compatible even when a particular channel does not yet have a dedicated field for that relationship.
Aiper pool cleaner ├─ compatible filter ├─ replacement part ├─ alternative model └─ accessory
For years, many ecommerce sites have represented these connections only through visual merchandising blocks such as “Frequently bought together,” “Compatible accessories” or “You may also like.”
AI retrieval gives businesses a reason to model the relationship itself, independently of how any one channel chooses to ingest it.
Connect important product knowledge to the product in your commercial model first. Then expose the subset each channel can currently consume.
This was one of the most useful parts of our research because it shows what an AI-readiness audit looks like in practice.
Clark Rubber is a Core dna customer, and this test was not a pass/fail exercise on the website. It was part of the kind of ongoing monitoring we do with clients: test how new discovery interfaces are using the experience today, identify where the next retrieval opportunity sits, and then work with the team on how the architecture should evolve.
In a logged-out ChatGPT session, we asked:
“I’m in Richmond, Victoria and need a pool robot under $1k by tomorrow, where can I get one, pickup or delivery?”
ChatGPT found several products and merchants. Importantly, it surfaced Clark Rubber and an Aiper Scuba S1.
So the basic discovery layer worked.
We then moved from discovery to a narrower retrieval question: could ChatGPT get to the store-level availability that Clark Rubber exposes through its Find in Store experience?
We tested the exact Aiper Scuba S1 product page directly.
ChatGPT could read the product itself, including the current $799 sale price, product details, specifications, reviews and the fact that it was marked “Exclusively In-Store.”
But it could not retrieve the individual store availability.
What it could see was the customer-facing prompt to use Find in store / Click & Collect. The actual store-by-store inventory did not appear in the readable product-page content available to the retrieval path used in this test.

A human shopper, however, can click Find in Store and see store-level statuses such as In stock and Low stock.

That gives us three distinct states:
| Layer | What happened |
|---|---|
| The business knows | Clark Rubber has store-level inventory information. |
| The human can retrieve it | Product page → Find in Store → local inventory. |
| The AI retrieval path in our test | Product ✓ Retailer ✓ Price ✓ Product details ✓ Store-level stock not yet exposed through this retrieval path |
The important point is not that the website “failed.”
The commerce platform already has the data and the customer experience already makes it available to a human. The audit simply identified the next layer of work: making that same operational information easier for AI retrieval systems to reach.
There is an important technical nuance here. ChatGPT itself described the limitation as the store inventory being loaded through the interactive Find in Store experience rather than exposed in the product content it could access. We should treat that as evidence of an interaction-dependent retrieval gap, not as proof that every crawler or browser-capable AI would fail in the same way.
This is now part of the next phase of our work with the Clark Rubber team: evaluating the cleanest way to expose store-level availability through a supported machine-readable retrieval path while preserving the customer experience already built around it.
That is also how we think about AI readiness more broadly. The interfaces are changing quickly, so the work is not a one-off implementation. It is an ongoing cycle of observe → test → identify the gap → improve the retrieval path → retest.
If your business already has the best answer to the customer's question, AI should have a reliable path to retrieve it too.
This is also why we distinguish retrieval from browser agents. A browser-capable AI may be able to execute the interaction itself. That is a different test, and one we will cover in our separate work on AI-side browsing and agentic commerce.
The Clark Rubber example exposes a broader problem with the way teams often audit digital commerce.
We tend to ask whether the customer can get to the information.
AI retrieval adds another question:
Through which technical surface is the information available?
| State | What it means | Example |
|---|---|---|
| Human-visible | A person can see the data after navigating the interface. | Store inventory appears after clicking Find in Store. |
| Rendered-page accessible | The information exists after scripts or interactions update the page. | Availability is injected client-side. |
| Crawler-accessible | A retrieval system can fetch and parse the information reliably. | Important product facts are present in crawlable HTML. |
| Structured | The information is expressed as explicit fields rather than inferred from prose. | SKU, availability, price, location ID. |
| Channel-ingested | The information is supplied through a feed or data source the destination actually supports. | Merchant Center product or local inventory data. |
| Integrated for retrieval | A live service or API is explicitly connected to the system answering the query. | A supported integration can resolve inventory by SKU and location. |
These states are not mutually exclusive, and not every product fact needs every layer.
But the distinction is useful when an AI answer is incomplete.
If the information appears to a human but not to the retrieval system, you may have an accessibility problem.
If the crawler can see it but the meaning is ambiguous, you may have a structure problem.
If the data is structured but stale, you have a freshness problem.
If a live API contains the answer but the AI system has no supported path to call it, you have an integration problem.

This is a much more useful way to think about AI retrieval readiness than simply checking whether Product schema exists or whether your platform has an API.
For years, ecommerce teams could maintain relatively clear boundaries between marketing content and operational systems.
Content belonged to marketing. Price belonged to commerce or ERP. Inventory belonged to inventory systems. Delivery belonged to OMS and fulfilment.
A conversational query collapses those boundaries:
“Find me this product under $1,000, in stock near Richmond, that I can collect tomorrow.”
Every part of that request can influence the answer.
OpenAI's current stable discovery feed makes price and availability required product fields. It also supports shipping data. Its documentation is explicit that merchants should update current price and availability when they change; uploaded sale-window or availability dates do not automatically schedule those commercial state changes.
Google goes further for local retail. Its local inventory data model can associate the same product ID with a specific store code, store-level availability, optional quantity and store-specific price. It also supports pickup information such as pickup method and pickup SLA.
This distinction matters because “available online” and “available at the store near me” are not the same commercial fact.
And all of these fields are volatile.
10:03 Store inventory changes 10:04 Inventory system knows 10:05 Ecommerce platform knows 10:06 Website knows ? Merchant feed updates ? AI retrieval source refreshes
The gap between those timestamps is now commercially meaningful.
Every commerce team should be able to answer:
What is our product-data freshness SLA for price, availability and fulfilment?
The answer will differ by business. A nightly export may be perfectly adequate for relatively stable products. It may be completely inadequate for a retailer with local inventory, flash promotions and same-day pickup.
Pay particular attention to low inventory, store inventory, fast-moving promotional pricing, limited releases, spare parts, B2B availability, delivery cut-off times and pickup promises.
OpenAI's current file-upload route itself is snapshot-based and recommends publishing a complete catalogue at least daily. That makes it useful for broad product discovery, but it also demonstrates why highly volatile local commerce information may need a different retrieval path.
Before going further, it is worth defining a term we use throughout this article.
By canonical commercial model, we mean the authoritative logical view of a product and the commercial facts needed to sell it: identity, attributes, variants, price, availability, relationships and fulfilment.
“Canonical” does not mean all of that information must physically live in one database.
In enterprise commerce, different systems often remain authoritative for different kinds of truth:
PIM → product specifications ERP → inventory Pricing → price and customer rules OMS → fulfilment CMS → editorial content DAM → media CRM → customer context
The architectural requirement is that those systems can resolve into one consistent commercial answer when a downstream channel needs it.

For a given product and customer context, that model should be able to resolve:
- What is the product?
- Which variant?
- Which market?
- What is the applicable price?
- Is it available?
- Where is it available?
- What can fulfil it?
- What does it work with?
- What policies apply?
This is particularly important for organizations with multiple brands, regions, dealer networks, franchises or B2B pricing. Copying the catalogue for each destination may appear easier in the short term, but every copy becomes another place where identity, price and availability can drift.
Canonical means authoritative and reconcilable, not necessarily centralized.
One of our engineers raised an important practical question while we were working through this research.
Many modern ecommerce sites depend heavily on JavaScript. Product variants update without a page reload. Store availability appears in a modal. Pricing, personalization and delivery estimates may all be resolved client-side.
So does making a catalogue easier for AI mean rebuilding the website without JavaScript?
No.
This article is about retrieval: whether an AI system can reliably obtain the commercial information it needs to answer a shopper's question.
For retrieval, the more useful principle is:
Do not make the JavaScript interface the only path to commercial truth.
JavaScript itself is not the problem
Google can execute JavaScript, and its own documentation explicitly supports JavaScript-powered sites. But Google also recommends putting Product structured data in the initial HTML for the best shopping results and warns that dynamically generated Product markup can make Shopping crawls less frequent and less reliable, particularly for fast-changing information such as price and availability.
Google also notes that not every bot can execute JavaScript.
The conclusion is not “remove JavaScript.”
It is: make sure the commercial facts required for retrieval also exist through a surface the relevant system can reliably consume.
1. Initial HTML and structured data
For stable, public product facts, the web page is still an important retrieval surface.
Identity, variant information, price, availability and other public facts should be represented in a way crawlers can reliably interpret. Depending on the architecture, that may involve server-side rendering, pre-rendering or ensuring the relevant structured data is present in the initial HTML.
The frontend can remain interactive. The important thing is that the commercial truth does not exist only inside the interaction.
2. Merchant feeds
A merchant feed is useful for distributing a normalized catalogue at scale.
OpenAI's current stable file-upload route accepts a full catalogue snapshot via SFTP and recommends publishing a complete snapshot at least daily. It is a real product-discovery integration, not a theoretical “LLM file.”
But there is an important geographic caveat.
OpenAI's documentation currently says the standard OpenAI-format upload targets the US. Additional markets require OpenAI to confirm the integration and supported countries and currencies.
That means global merchants should treat this as an emerging channel, not a universal feed destination today. Coverage may broaden over time, but architecture should be based on documented availability rather than an assumed rollout timeline.
3. APIs are useful, but API-accessible does not mean AI-accessible
This is one of the most important distinctions in the article.
A company can have an excellent API and still have poor AI retrieval.
An API being technically available means another system could call it. It does not mean ChatGPT, Google or another AI system knows that the endpoint exists, has permission to use it, understands the contract or is wired into it during a customer query.
API-accessible ≠ AI-accessible
API accessibility is a system capability. AI accessibility is an integration capability.
Core dna itself is a useful example. The platform exposes roughly 300 API endpoints. That does not automatically make all of those capabilities available to an AI model. We separately expose a governed MCP layer with around 80 tools so an AI system can discover and call specific capabilities intentionally.
The same logic applies to ecommerce inventory.
A retailer may already have:
GET /inventory?sku=123&store=456
But unless that endpoint is used to generate a supported feed, exposed through an AI-compatible tool layer, or integrated into the system answering the query, the existence of the API alone does not solve the retrieval problem.
This is exactly why the Clark Rubber example matters. The business has the inventory data. The technical question is how that data reaches the AI retrieval surface being used by the customer.
4. llms.txt is optional infrastructure, not a commerce database
llms.txt is an emerging community proposal for giving AI systems a concise, LLM-friendly map of a website and pointing them toward useful pages or Markdown representations.
We think it is worth considering. It is lightweight, and there is little downside to providing a clean machine-oriented index of important content. We have a separate llms.txt guide that goes deeper into what it does, what it does not do and where it fits in an AI-search strategy.
But it should not be confused with Merchant Center, an OpenAI product feed or a live inventory service.
For ecommerce, it may be useful for pointing machines toward:
- product documentation;
- important category resources;
- technical manuals;
- shipping and returns policies;
- machine-readable versions of key content.
It is not the right source of truth for thousands of SKUs with changing price and store inventory.
llms.txt can help an AI find useful catalogue knowledge. It should not become the catalogue.
Choose the retrieval surface based on the data
| Type of commercial data | Examples | Useful retrieval surfaces |
|---|---|---|
| Relatively stable | Name, brand, material, dimensions, specifications | Initial HTML, structured data, merchant feed |
| Moderately dynamic | Price, promotion, online availability | Structured data + frequently refreshed feed |
| Highly dynamic | Store stock, pickup availability, delivery promise | Local inventory feed and/or an explicitly integrated real-time service |
| Knowledge / documentation | Manuals, FAQs, policies, technical guidance | Public pages, structured fields, Markdown, optional llms.txt links |
This gives us a cleaner rule:
Keep the human experience as rich as it needs to be. Just make sure important commercial information also has a supported machine-readable retrieval path.
A note on what comes next
Retrieval is only the first layer.
An AI answering “Where is this product in stock?” is a retrieval problem. An AI opening the site, navigating the interface and taking an action for the customer is a browser-agent or agentic-commerce problem.
Those technologies are still developing, and adoption will depend on more than technical capability. Customers and businesses will have to become comfortable with issues such as authentication, permissions, personal data, transaction approval and trust.
We are treating that as a separate research topic because the architecture and adoption questions are substantially different from the retrieval problem covered here.
Finding 11: Do not build a separate ChatGPT catalogue
A common reaction to a new channel is to create a new feed, a new spreadsheet and eventually a new source of truth.
Google catalogue ChatGPT catalogue Marketplace catalogue Website catalogue AI-search catalogue
That recreates the problem every time a new interface appears.
The current specifications already show why channel-specific modelling is fragile.
OpenAI's stable discovery feed uses concepts such as group_id, listing_has_variations and variant_dict for variant grouping.
Google uses its own Merchant Center product model, including item_group_id, standard variant attributes and newer conversational fields such as variant_option, Q&A and related products.
Those schemas are not identical, and they do not need to be.
The architectural goal is:
┌─ Website
├─ Google
├─ ChatGPT
├─ Marketplaces
CANONICAL COMMERCIAL ────┼─ AI search
MODEL ├─ Customer service
├─ Sales tools
└─ Schema / feedsThe canonical commercial model contains the business truth.
Each destination gets a mapping of the fields it currently understands.
If Google introduces a new conversational attribute, you map into it. If OpenAI changes its discovery contract, you update the output. The underlying product should not have to be remodelled every time a channel evolves.
One commercial model. Many channel mappings.
Change the mapping. Do not rebuild the product.
Finding 12: Structured does not automatically mean trusted
There is another layer that should remain separate from product structure.
Structure helps answer:
What is being claimed about this product?
Identity and provenance help answer:
Who is making the claim, and which source should the retrieval system treat as authoritative?
AI retrieval needs both.
Three different jobs
Structure → what is being claimed
Provenance → who is making the claim
Governance → which source wins when systems disagree
A verified merchant with contradictory prices, variants and inventory still has a commercial-data problem.
A beautifully structured catalogue from an unreliable or stale source has a trust problem.
This is why an AI-ready retrieval architecture also needs governance around:
- merchant and source identity;
- first-party domains;
- authenticated feeds where required;
- stable identifiers;
- authoritative source systems;
- review provenance;
- audit history;
- ownership and approval of commercial data.
This matters especially when multiple systems can publish overlapping information. If the website says one price, the merchant feed says another and a marketplace says a third, the problem is not that the AI needs a better prompt. The business needs a clearer commercial source of truth and a reliable update path.
Structure does not replace trust. Trust does not replace structure.
When evaluating an eCommerce platform for AI shopping, the least useful question may be:
“Does it have AI?”
A better question is:
“Can the platform represent our commercial truth cleanly and expose the right parts through the retrieval surfaces each channel actually supports?”
| Platform capability | Why it matters for retrieval |
|---|---|
| Stable product and variant IDs | Allows the same product to be reconciled across storefronts, feeds and AI-search surfaces. |
| Typed product attributes | Gives machines explicit facts rather than forcing inference from copy. |
| Strong variant modelling | Connects answers to the actual SKU a customer can purchase. |
| Product relationships | Lets the business model compatibility, required parts, substitutes and bundles even when a channel supports only a subset. |
| Regional catalogue rules | Prevents irrelevant or unavailable products being exposed in the wrong market. |
| Dynamic / contextual pricing | Provides the correct commercial answer for market, customer or contract. |
| Store and local inventory | Creates the underlying truth needed for “where can I get it?” queries. |
| Fresh availability | Reduces stale product recommendations and failed pickup promises. |
| Fulfilment integration | Supports delivery-date, shipping-cost and pickup constraints. |
| Structured-data generation | Makes important public product facts easier for crawlers to interpret. |
| Feed generation | Lets Google, OpenAI and other destinations become outputs rather than independent catalogues. |
| Incremental / frequent updates | Keeps volatile commercial data closer to reality. |
| APIs | Provide a programmable source for live or contextual data, but still require an explicit integration before an AI system can use them. |
| Governance and audit history | Makes changes to commercial truth accountable and traceable. |
| PIM / ERP / OMS integration | Allows existing systems to remain authoritative where appropriate. |
There is one question that cuts through most platform discussions:
If a major new AI-shopping or AI-search interface launched tomorrow, could you expose the appropriate version of your catalogue to it without remodelling your products?
If yes, you are in a strong architectural position.
If no, the problem is bigger than ChatGPT optimization.
See what an AI-ready commerce architecture looks like in practice.
A practical AI product-catalogue audit
You do not need to start with the whole catalogue.
Start small, but make the sample difficult
Take 25 to 50 representative products and deliberately include the difficult ones: a simple product, a product with many variants, a discounted product, a low-stock product, an item with local inventory, a product with accessories, a replacement part, a configurable product, a B2B product if relevant, and something with complex fulfilment.
1. Identity
Does every purchasable product or variant have a stable ID? Do SKU, GTIN and MPN exist where appropriate? Can the same item be reconciled across your storefront, Merchant Center, marketplace and AI feed?
2. Attributes
Which buying characteristics are explicit structured fields? Which still live only in descriptions, CMS blocks, images or PDFs?
3. Variants
Does every relevant variant have its own identity, price, inventory, availability, image and attributes?
4. Relationships
Can the business model accessories, required parts, substitutes, bundles, compatible products and replacement models, regardless of whether every destination supports those fields yet?
5. Price
Where does price originate? Can it vary correctly by market, currency, account, contract or customer segment? How quickly do changes propagate?
6. Inventory
Is inventory global, regional, warehouse-specific or store-specific? Where is that data exposed today: page, feed, local-inventory service or API?
7. Fulfilment
Can your systems resolve “Can I get this tomorrow?”, “Can I collect it today?” and “Which store has it?” If the business knows the answer, is there a supported retrieval surface that can expose it?
8. Public product page
Can machines retrieve the important facts directly? Check structured data, canonical URLs, rendering, bot protection, dynamically loaded UI, PDFs and information hidden behind interactions.
9. Feeds and integrations
Are channel feeds generated from authoritative data? Which platforms can ingest them today? Where you have APIs, are they actually connected to an AI-consumable integration or are they simply technically available?
10. Governance
Who owns each piece of commercial truth? Is there approval, audit history, version control and a clear update path when price, inventory or product information changes?
Once you know where the likely gaps are, the retrieval test below lets you see how those gaps show up in a real AI answer.
Test your own product catalogue with AI
You do not need a complex benchmarking platform to start testing whether AI can actually understand your products.
The most useful approach is to ask the kinds of questions a real customer would ask, then compare the answer with the commercial truth in your own systems.
This is not a ranking test. A product failing to appear in one answer does not prove your catalogue is broken. Treat it as a blind-spot test for discovery, product structure, price, inventory, fulfilment and relationships.
Here are three examples from the full test:
Unbranded discovery
“I need a [product category] under [budget] for [specific use case]. What would you recommend and where can I buy it?”
Look for: whether your product appears without naming the brand, which alternatives appear and which sources are cited.
Local inventory
“I’m in [suburb/postcode]. Where can I buy [product] in stock today?”
Look for: whether the AI can resolve store-level availability or stops at a generic “check store availability” answer.
The stress test
“I’m in [location]. I need [product category] under [budget], with [attribute], compatible with [existing product], available for pickup by [date]. What are my best options?”
Look for where the answer breaks when identity, attributes, compatibility, price, inventory, location and fulfilment all have to resolve together.
We have turned the full framework into a downloadable AI Product Catalogue Retrieval Test with all 15 copy-and-paste prompts, guidance on what to check after each one, a results scorecard and a simple diagnostic for identifying where retrieval breaks.
Run the same prompts again after catalogue, feed or integration changes. The before-and-after comparison is much more useful than a one-off AI result.
One useful way to frame the work is as five levels rather than a binary “AI ready / not AI ready.”
| Level | State | What it looks like |
|---|---|---|
| 1 | Visible | Products have public, crawlable pages and stable URLs. |
| 2 | Machine-readable | Important product facts, variants and offers are explicitly structured. |
| 3 | Syndicated | Google, marketplaces and supported AI feeds are generated from authoritative catalogue data. |
| 4 | Context-aware and fresh | Price, inventory and fulfilment can be resolved at the right market, variant and location level with an appropriate update cadence. |
| 5 | Retrieval-ready | Important commercial facts are exposed through supported machine surfaces and tested against real customer queries to identify blind spots. |
The important thing is that the earlier levels remain necessary as you move upward.
A feed does not fix poor variant modelling. A real-time API does not help an AI system unless there is an integration path to it. Perfectly structured data does not help if it is stale. And appearing once in an AI answer does not prove the entire catalogue is retrieval-ready.
AI retrieval readiness is a stack of capabilities, not one optimization tactic.
What to do in the next 30 days
Most organizations do not need a six-month AI-commerce transformation before they can learn something useful.
Start with a focused month-long audit.
Week 1: Test what AI can currently see
- Choose 10–20 high-value products.
- Run unbranded and branded conversational shopping queries.
- Add constraints: size, location, compatibility, budget, availability, delivery date.
- Record what the AI gets right, what it misses and which source it uses.
Week 2: Trace the commercial truth
- Map which system owns identity, attributes, price, inventory, fulfilment and content.
- Identify important facts buried in prose, PDFs or UI interactions.
- Measure update latency for price and stock.
Week 3: Fix the highest-value gaps
- Normalize product and variant IDs.
- Move key buying attributes into structured fields.
- Improve Product / Offer structured data.
- Fix feed conflicts and stale availability.
- Expose one high-value operational data gap through an appropriate machine interface.
Week 4: Retest
- Repeat the same queries.
- Compare retrieval quality.
- Document remaining gaps.
- Prioritize the catalogue, API and governance work that has the highest commercial value.
The 30-day goal
Not “rank in ChatGPT.” Understand whether your digital commerce stack can answer the questions customers are increasingly delegating to AI.
Traditional product-feed work is not disappearing.
Titles matter. Descriptions matter. Images matter. Taxonomy matters. Identifiers matter.
But the research points to a larger shift.
OpenAI's current stable discovery feed requires current product identity, price and availability and supports richer variant, attribute, shipping, returns and review data.
Google is adding conversational product attributes for Q&A, documentation and product relationships, alongside AI-specific reporting that shows which terms, attributes and intents appear in conversational shopping queries.
Perplexity explicitly connects richer merchant data with recommendation relevance.
And our ongoing AI-readiness work with Clark Rubber gave us a practical example of where the next improvement sits: the business already has the store-level information a shopper needs, and our retrieval audit identified an opportunity to make that information easier for AI systems to access.
That changes the question.
It is no longer only:
Is our product feed optimized?
It becomes:
Can the AI systems our customers use retrieve an accurate, structured and current representation of what we sell?
Do not remodel your business around ChatGPT.
Do not remodel it around Google.
And do not create another independent catalogue every time a new AI-search surface appears.
Model the commercial truth properly, then make each retrieval channel an output.
This article has deliberately focused on retrieval. The next layer is different: AI-side browsing, browser agents and increasingly merchant-owned commerce agents that can search a catalogue, assemble a cart, support checkout and answer post-purchase questions.
Anthropic's new Claude commerce blueprint is a useful example of that direction. It is designed for merchants to build agents inside their own experience, connected to their own systems, rather than relying only on a third-party AI to retrieve product information from the outside. Anthropic explicitly positions Claude as the intelligence layer rather than the storefront or checkout.
Industry reaction: the AI layer is not the hard part
The launch also triggered a useful reality check from people who have spent years inside commerce infrastructure. Kelly Goetsch, President of Pipe17, argued that the difficult part is not adding a conversational layer. It is connecting that layer to the operational complexity underneath it: catalogue logic, inventory, pricing, fulfilment, supplier systems, ERP, EDI and the business rules that have accumulated over years.
We think that criticism is important to include because it tempers the launch without dismissing it.

The AI layer is getting easier to build. The commercial infrastructure underneath it is becoming the differentiator.
That distinction matters because agentic commerce introduces a different architecture and a different adoption problem: tool access, permissions, identity, transaction authority, personal data, payment handoff, safety controls and customer trust.
We are working on a separate article to tackle that next phase properly rather than collapsing it into the retrieval problem here.
The broader architecture and governance questions are covered in our Enterprise AI Commerce Readiness Guide.
Merchants specifically evaluating paid placement can go deeper in ChatGPT Ads and eCommerce: What Merchants Should Actually Do Now.
Summarize with