What problem does it solve?
An online catalog is only as findable as its data. Shoppers filter by size, color and material, search engines index titles, descriptions and alt text, and marketplaces such as Amazon use attributes like color to index products in their own search. Yet most product records arrive thin or inconsistent: a supplier spreadsheet with a cryptic name, a few bullet points and a photo. Retailers with hundreds of thousands or millions of items cannot write and check every record by hand, so attributes stay empty, descriptions stay copied from the manufacturer, and items that are in stock are never found.
Walmart describes the stakes plainly: the quality of catalog data affects nearly everything it does, from helping customers find products to sorting inventory and delivering orders. Small sellers feel the same problem from the other side. Writing a good listing takes time they do not have, which is why the large marketplaces now offer to draft it for them from a photo, a few words or an existing web page.
How does it work?
- Collect the inputs. Supplier feeds, existing descriptions, product images, the category and the attribute specification for that category (allowed values, units, required fields).
- Extract attributes. A model reads the text and the images and proposes a value for each attribute in the specification, with the source it used.
- Write the content. A model drafts the title, description, bullet points and image alt text in the house style, using only the extracted attributes and approved claims, and adds the keywords shoppers use in search.
- Check before publishing. A second model or rule set checks each value and each draft: does the attribute match the image, is the unit valid, does the description claim anything the data does not support? Values above a set accuracy threshold are published; the rest go to a human reviewer, whose decisions become new training and test data.
- Measure and repeat. Search conversion, returns for "not as described" and reviewer corrections show which categories and attributes need better prompts or a human.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Mainstream
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Interactions handled | Not pooled | 100 million to 850 million | 2 | 2 organization |
| Users served | Not pooled | 900,000 to 10 million | 2 | 2 organization |
| Quality score uplift | Too few to pool | 40% | 1 | 1 organization |
Value drivers: Revenue growth, Employee productivity, Customer experience, Speed and cycle time.
Indicative value
An online retailer with 200,000 active product listings
USD 300,000 to USD 2.5 million
Content production cost avoided per year
How this is calculated
Formula: listings * shareTouched * minutesSaved / 60 * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Active product listings listings, listings | 200,000 | 200,000 | The reference retailer. |
| Share of listings written or repaired per year shareTouched, fraction of listings | 0.3 | 0.6 | Editorial assumption covering new items and repairs of thin or inconsistent records. Replace with your own assortment change rate. |
| Content work saved per listing minutesSaved, minutes per listing | 10 | 25 | Editorial assumption. A seller quoted by Amazon in May 2025 (in the Amazon evidence record for this page) said listings used to take an hour and that the AI content is now generated in under 15 minutes; that 15 minutes is generation time, not total listing time. This range is more conservative because a retailer's content team already works faster than a small seller and still reviews each item. |
| Fully loaded cost of a content specialist hourlyCost, USD per hour | 30 | 50 | Editorial assumption, replace with your own cost or agency rate. |
What it leaves out: Counts only the writing and data entry time saved. It leaves out the running cost of the models, the human review that stays in place, and the revenue effect of better findability, which Etsy reports as 3% more conversions for its sellers (alongside 5% more visits from search engines) from improved alt text, but which depends on the catalog.
Who already uses it?
4 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Amazon
United States · Retail and ecommerce · 2025
Since the end of 2023 Amazon lets independent sellers create a product listing from a few words or a single image, and since March 2024 also from the URL of their own web page: generative AI on Amazon Bedrock drafts the title, bullet points, description and attributes, and the seller submits the draft, with Amazon encouraging a review first. Bulk creation from a spreadsheet followed, and Enhance My Listing, which Amazon said in May 2025 had begun rolling out in the US, suggests updates to existing listings based on shopping behaviour. In May 2025 Amazon reported that sellers accept the AI generated content with little to no edits about 90% of the time. The same post describes Amazon using generative AI to personalize product recommendation categories and product descriptions shown to customers on the website and in the shopping app, based on a customer's shopping activity; no outcome number is given for that side of the work.
- Users served: at least 900,000, selling partners who have used the listing tools, by May 2025
"Now, more than 900,000 Amazon selling partners have embraced these tools, with sellers accepting AI-generated content with little to no edits approximately 90% of the time."
Claimed by: organization - Quality score uplift: 40%, overall listing quality of listings created with the tools, reported by May 2025
"When sellers use our Gen AI tools to create listings, they see a 40% increase in overall listing quality, helping them create content that enhances customer engagement and boosts sales potential."
Claimed by: organization
eBay
United States · Retail and ecommerce · 2025
eBay's magical listing tool writes item descriptions from known product attributes and, from a seller's photo, suggests the category and item specifics for the seller to review and approve. The description feature reached all sellers in eBay's top five markets by the end of 2023; a bulk version creates drafts from batches of photos, and in 2025 a simplified mobile flow starting from photos was rolled out in the US, UK and Germany. Early UK tests showed half as many steps to list.
- Users served: at least 10 million, sellers worldwide who have used any eBay AI feature (eBay AI overall, not only listing generation), by April 2025
"Over 10 million sellers worldwide have already used eBay’s AI features, creating well over 100 million listings using AI and generating billions in gross merchandise volume (GMV)."
Claimed by: organization - Interactions handled: at least 100 million, listings created with any eBay AI feature (eBay AI overall, not only listing generation), by April 2025
"Over 10 million sellers worldwide have already used eBay’s AI features, creating well over 100 million listings using AI and generating billions in gross merchandise volume (GMV)."
Claimed by: organization
Walmart
United States · Retail and ecommerce · 2024
Walmart uses several large language models to extract product attributes such as color, size and material from item descriptions and images and to create or improve catalog data. One model extracts the values and a second model, tuned on human validated labels, checks them; values for attributes above an accuracy threshold go into the catalog, and the rest are checked by the quality model, with specialists validating samples. Walmart's chief executive told investors in August 2024 that the work covered more than 850 million pieces of catalog data, and estimated that without generative AI the same work would have needed nearly 100 times the current headcount to finish in the same time.
- Interactions handled: at least 850 million, catalog data points created or improved, reported August 2024
"We've used multiple large language models to accurately create or improve over 850 million pieces of data in the catalog."
Claimed by: organization
Etsy
United States · Retail and ecommerce · 2025
Etsy uses Gemini models with BigQuery and Dataflow to enrich data about more than 130 million items listed by more than 5 million sellers: classifying items, spotting items linked to emerging trends and generating better image alt text for listings. Etsy's engineering lead for search says the improved alt text increased visits from search engines by 5% and conversions by 3% for sellers.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- An attribute specification per category, with allowed values and units
- A human validated sample of products per category to benchmark accuracy
- Brand style guide and a list of claims that may and may not be made
- Product images and supplier data linked to each item
Systems to integrate
- Product information management (PIM) or catalog system
- Digital asset management for product images
- Ecommerce platform or marketplace listing API
- Search and analytics for conversion and search performance
Complexity: Medium
Drafting text is easy; getting reliable attributes at catalog scale is the work. It needs an attribute specification per category, a labelled benchmark set, a quality check that decides what is published automatically, and a connection to the product information system.
- 1
Write the specification first
For each category, list the attributes that matter for search and filters, their allowed values and units, and which ones are safety or compliance relevant. The model is only as good as this list.
- 2
Build a benchmark before a prompt
Have specialists label a few hundred items per large category, keep them out of any tuning data, and measure precision and recall per attribute for every model or prompt change.
- 3
Separate writing from checking
Use one step to extract and write and another to verify, as Walmart describes. Publish automatically only the attributes whose measured accuracy clears your threshold, and send the rest through a further check or to reviewers.
- 4
Ground descriptions in the data
Generate descriptions from the verified attributes and approved claims only, never from the model's general knowledge, so a description cannot promise a feature the item lacks.
- 5
Put people where the risk is
Keep human review for safety relevant attributes (allergens, age ratings, electrical ratings), regulated categories and low confidence items, and sample the automatic output every week.
- 6
Measure on the storefront
Track search conversion, zero result searches, filter usage and returns for "not as described" per category, so enrichment is judged by shoppers, not by word count.
Guardrails
- Descriptions generated only from verified attributes and an approved claims list
- Automatic publication only above a measured accuracy threshold per attribute
- Human review for safety, legal and regulated product attributes
- Filters that block model refusals, placeholder text and competitor brand names from publication
- Versioned prompts and a benchmark run before every change
KPIs to instrument
- Attribute precision and recall per category against the human labelled benchmark
- Share of AI output published without edits, and the edit rate by reviewers
- Attribute fill rate for the fields shoppers filter on
- Search conversion and zero result searches before and after, per category
- Returns with the reason "not as described"
Human in the loop
Content and category specialists own the attribute specifications, label the benchmark sets, review low confidence and safety relevant values, and sample automatically published content. On a marketplace the seller submits each AI draft: Amazon encourages sellers to review drafts before they submit them, and eBay's listing flow asks the seller to review and approve the suggestions.
Common failure modes
- Confident but wrong attributes
- A model fills a color, size or material that the image contradicts, and the item is returned. Measure accuracy per attribute and only automate what clears the threshold.
- Invented features
- A fluent description promises a feature the product does not have, which becomes a consumer protection problem. Generate only from verified data and approved claims.
- Unreviewed output goes live
- In January 2024 Amazon hosted listings whose titles were model refusal messages. Block refusals and placeholder text automatically and keep a human submit step.
- Thin pages at scale
- Thousands of near identical generated pages can be treated as spam by search engines. Write for shoppers with real product facts, not for keyword variants.
What are the risks and rules?
EU AI Act
Depends on design
Writing product content and extracting catalog attributes is not an Annex III use and makes no decisions about people. When a retailer uses a third party generator, the use is minimal risk for the retailer: the Article 50(2) duty to mark generated text in a machine readable way falls on the provider of that system. When a retailer builds and operates its own generating system and puts it into service under its own name, it is the provider and must mark the output, unless the exception for systems that only assist standard editing or do not substantially alter the input applies. Article 50(4) covers text published to inform the public on matters of public interest, not product listings. Consumer protection law applies to what the listing says in every case.
Rules that apply
Guidance
- Google Search's guidance about AI generated content (Google Search Central, Global). Google rewards helpful content however it is produced, and treats automation used mainly to manipulate rankings as spam.
- Spam policies for Google web search, scaled content abuse (Google Search Central, Global). Lists pages generated at scale with little value for users, including automated transformations such as synonymizing and translating, as scaled content abuse.
- Final rule banning fake reviews and testimonials (Federal Trade Commission, North America). Prohibits fake reviews and testimonials, including AI generated ones; product content generation must not extend to reviews.
- Unfair Commercial Practices Directive (European Union, Europe). Misleading product information is an unfair commercial practice, whoever or whatever wrote it.
Controls to put in place
- A named owner per category for the attribute specification and the approved claims list
- Accuracy benchmarks per attribute with thresholds for automatic publication
- An audit trail that records which model, prompt version and reviewer produced each published value
- Alt text that describes the image for people using screen readers, not only for search engines
- A takedown route for content reported as wrong by customers or sellers
When it went wrong elsewhere
- AI refusal messages published as Amazon product titles. Amazon hosted listings whose titles were model refusal messages ("I'm sorry but I cannot fulfill this request"); Amazon removed them and said it was improving its review systems. The date comes from the incident ID (2024-01-12). It shows what happens when generated content is published without review.
Frequently asked questions
- Can AI write product descriptions without a human checking them?
- For extracted attributes with measured accuracy, yes, and that is how Walmart describes its catalog work: a second model checks the first, and values for attributes above an accuracy threshold go straight into the catalog. Generated descriptions are a different matter: on Amazon and eBay the seller submits each draft, and Amazon reported in 2025 that sellers accept AI content with little or no edits about 90% of the time.
- Does AI generated product content hurt SEO?
- Not by itself. Google says it rewards helpful content however it is produced, but treats pages generated at scale mainly to manipulate rankings as spam. Etsy reports that better alt text generated with Gemini increased visits from search engines by 5% and conversions by 3% for its sellers.
- What is harder, the text or the attributes?
- The attributes. Fluent descriptions are easy to generate; correct size, color, material and safety data across millions of items need a specification, a labelled benchmark and a quality check per attribute.
How to cite this page
Blits.ai AI Use Case Library, "AI product content and catalog enrichment for online retail", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/product-content-and-catalog-enrichment. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published