Recommendation engines: how recommender systems work and what’s changing for eCommerce
AI is changing how businesses operate at every level, and recommendation engines are one of the first places customers feel it. Our CEO, Sam Saltis, spent a week at RecSys 2026 hearing how Google, YouTube, Amazon, Spotify and Target build theirs and what they're changing next.
A recommendation engine, also called a recommender system, is the software that decides what each customer sees, from product rows and search results to the answer an AI shopping assistant gives. Netflix has said about 80% of what its members watch comes from recommendations. McKinsey has estimated Amazon's figure at 35% of purchases.
On this page:
What is a recommendation engine (or recommender system)?
A recommendation engine is software that predicts which items each person is most likely to want, and chooses what to show them. Recommendation engine, recommender system and recommendation system all mean the same thing. An online recommendation engine is one that runs on a website or app.
Most shoppers meet one many times a visit without noticing. Anywhere a store chooses what to show one person, instead of showing everyone the same thing, a recommender is usually making the call.
Where customers meet one
- Product pages: similar items, frequently bought together, complete the look.
- Home and category pages: “recommended for you,” trending products, personalized sorting.
- Search: which results appear first, and which products appear for a vague query.
- Email, SMS and push notifications: which products, content or offers each person receives, and when.
- Deals and promotions: which coupon or discount each customer is offered.
- Content feeds: articles, videos and posts, ordered for each person.
- AI assistants: the products, plans and baskets a chat assistant suggests.
Target gave a simple example at RecSys 2026, the main research and industry conference for this field. A guest with young children may see baby and children’s categories when they open the website, while a guest who mainly buys groceries sees grocery. The same system picks which deals and which marketing content each guest sees.
Recommendation and personalization
Personalization is the broader practice of tailoring an experience to a person. A recommendation engine is the system underneath that picks and orders what they see. We cover the strategy side in eCommerce personalization, and how a personalization engine works behind the scenes in a conversation with our CTO.
Search is a recommender too
Search used to rank pages against the words in a query. It now also uses what it knows about the person asking. Justin Basilico, who led personalization research at Netflix for years and now works on Google Discover and Search, presented Google’s work on this at RecSys 2026. We cover what it means for SEO and AEO further down.
Site search is moving the same way. Personalized search ranking orders results using what a store already knows about the shopper, such as their location, membership or B2B contract.
How do recommender systems work?
You don’t need the mathematics to make good decisions about a recommender. You do need to know what it learns from, how it narrows millions of options to a handful, and where it tends to go wrong.
What it learns from
- Behavior: browsing, searches, clicks, add-to-cart, purchases and returns.
- Product data: category, brand, attributes, price, descriptions and images.
- Context: location, season, device and time of day.
- Stated preferences: what customers tell you directly, in forms, filters or conversation.
- Similar customers: what people with comparable behavior went on to choose.
How it matches people and products
Modern systems often represent customers and products as embeddings: positions in a shared space, where a customer sits close to the products they are likely to want. Target uses guest and item embeddings as one layer of its personalization stack. How a system learns those positions, from behavior, from product attributes or from both, is what separates the types of recommender system covered in the next section.
The funnel

A catalogue can hold millions of items, and a page has room for a handful. Most systems narrow the field in stages:
- Retrieve: fetch a few hundred plausible candidates quickly.
- Rank: score those candidates for this person, using richer and more expensive models.
- Re-rank: apply business rules such as stock, margin, promotions, variety and items that should never appear together.
The funnel is so well established that Amazon now uses it for a different job. Its team described choosing which AI model should power each shopping feature by treating the available models like a catalogue: filter on hard constraints such as hardware, speed and cost, score the rest quickly, test the most promising, then fully train only a shortlist.
What a mature stack looks like

Target shared the layers behind its recommendations, a useful reference for what a mature setup includes:
- Guest and item embeddings.
- Affinity and propensity models for brand, category, attribute, style, price and deals, down to details such as color, size, fit and gluten-free.
- Sequential models that predict what someone is likely to want next from their history.
- Similarity models for similar items and frequently bought together.
- Retrieval, ranking and re-ranking.
Those layers power experiences such as “Recommended for you,” “Complete the look” and “Your top deals.”
Where recommender systems go wrong
- Cold start: new customers and new products have no history to learn from. Teams usually fall back on product attributes, popular or trending items, and asking new customers what they want. Google described its Discover feed as being in “eternal cold start,” because it focuses on fresh content. At Core dna, we’re working on a method called gradient superposition, which learns a new customer’s preferences from their first few interactions, in seconds. We explain how it works at the end of this guide.
- Noisy signals: a click can mean real interest, passing curiosity or an accident. Separating them is one of the hardest parts of the job.
- Position bias: items at the top get clicked partly because they are at the top, which can make a system overrate whatever it already shows.
- Feedback loops: a system trained on data its own recommendations shaped can narrow what people see over time. Research presented at RecSys has shown loops like this make behavior more uniform and recommendations less useful.
- Short-term targets: optimizing for the next click is easy to measure and can work against repeat purchase and loyalty.
Types of recommendation engines and the algorithms behind them
There are a handful of main types. Most production systems combine several of them, so the useful question about any system is which types it uses, and for what.

Collaborative filtering
Collaborative filtering recommends what similar people liked. It is the logic behind “customers who bought this also bought,” and it needs no understanding of the products, only behavior. It comes in two main forms: user-based (people like you bought this) and item-based (people who bought this also bought that). Matrix factorization, popularized by the Netflix Prize, is the classic algorithm. Its weakness is new products and new customers, which have no behavior to learn from.
Content-based filtering
Content-based filtering recommends items that resemble what someone already likes, by comparing product attributes such as category, brand, color, style and description. It can recommend a new product as soon as its attributes are filled in. It tends to recommend more of the same, and it is only as good as the product data behind it.

Hybrid recommender systems
Hybrid systems combine collaborative and content-based methods, for example by blending their scores, switching between them depending on how much data is available, or using one to fetch candidates and another to rank them. Most large production systems are hybrids.
Context-aware recommender systems
Context-aware systems adjust for the situation: location, season, device, time of day, or whether someone is shopping for themselves or for a gift. A grocery recommender that knows it is the week before a holiday is context-aware.
Knowledge-based recommender systems
Knowledge-based systems use rules and explicit requirements, such as compatibility, specifications, a budget or a B2B account’s approved catalogue. They suit complex or infrequent purchases, such as replacement parts, where there is little behavior to learn from.
Deep learning models
Neural networks learn from very large volumes of behavior. Two designs are common. Two-tower models learn embeddings for customers and items separately, so candidates can be fetched quickly. Sequential models predict what someone will want next from the order of their actions, which is how Target predicts a guest’s likely next item.
Generative recommendation
Generative recommendation is the newest type, and it takes two forms. Some systems predict the next item the way a language model predicts the next word; Spotify represents each item as a short learned code for this. Others use large language models to write the recommendation itself: a summary, a meal plan or a full basket. AI recommendation engines built on large language models fall into this group.
Recommendation engines in eCommerce
In eCommerce, a product recommendation engine decides which products each shopper sees, and in what order. It runs on almost every page of an eCommerce platform and in most channels, so its quality shows up across the whole store.
Where product recommendations appear
- Product pages: similar items, alternatives when something is out of stock, and frequently bought together.
- Cart and checkout: add-ons and accessories that complete the order.
- Search and category pages: the order of results, and what appears for a vague query.
- Homepage and landing pages: personalized sections and recently viewed items.
- Email, SMS and push: abandoned cart, back in stock, replenishment and new arrivals.
- AI assistants: baskets and plans assembled from a single request.
What a product recommendation engine needs
- Complete product data. Attributes, variants, compatibility and relationships between products, in structured fields. A single product catalog shared across storefronts keeps them consistent. Content-based and generative methods depend on these fields, and they are the main way to recommend new products that have no sales history.
- Current price and stock. Recommending something out of stock, or at the wrong price, costs the sale and the customer’s trust, so recommendations need live inventory data.
- Business rules. Margin, promotions, brand agreements and items that should never appear together, applied at the re-rank stage.
- Content alongside products. Buying guides, recipes, articles and videos can be recommended too. A content recommendation engine works best when it can draw on the same product data, so a guide and the products it mentions appear together. That is simplest when the CMS and the catalog run on one platform.
- Honest measurement. Live experiments and long-term outcomes such as repeat purchase and returns, covered later in this guide.
B2B recommendation engines
B2B buying works differently, and the recommendation engine has to reflect that. We cover the wider differences in B2B eCommerce personalization.
- Account-specific catalogues and pricing: a buyer should only see the products and prices their contract allows, so recommendations need to follow the same account rules as the B2B commerce platform.
- Reorders: many B2B orders repeat on a cycle, so predicting when a customer will run low is often more useful than suggesting something new.
- Compatibility: replacement parts, consumables and accessories need to fit the equipment the customer already has.
- Several people per account: the person browsing is often not the person approving the order.
- Less behavior to learn from: fewer, larger orders mean product relationships and knowledge-based rules often carry more of the load.
Questions to ask before choosing a recommendation engine
- What data does it need, and can it read our product attributes, variants, price and stock as they change?
- How does it handle new products and first-time visitors?
- Can we set business rules, and see when they are applied?
- What does it optimize for, and can we change that to repeat purchase or margin?
- How are results measured, and is there a live experiment with a control group?
- Can we see why a product was recommended?
- Does it work across the site, search, email and AI assistants, or only on product pages?
- Where does customer data go, and what is kept?
What is already live in 2026
Several changes discussed at RecSys 2026 are already in front of customers. For a store, the common thread is that an AI system now writes or assembles the recommendation, and it can only use product data it can read.
Shopping assistants that start from a goal

Customers increasingly ask for an outcome, and the assistant assembles the products.
- Amazon, whose shopping assistant is now called Alexa for Shopping, runs each conversational feature on its own specialized language model, chosen to keep answers fast, affordable and reliable.
- Target showed an on-site assistant that turns “help me plan meals for the week” into a five-day plan and a shopping list.
- Infobip demonstrated a WhatsApp shopping assistant that asks up to three clarifying questions, then returns a grouped basket. Asked for everything needed to make Neapolitan-style pizza, it came back with flour, yeast and mozzarella, each with a price.
AI-written summaries and notifications
Google Discover, the personalized feed in the Google app and Chrome, generates short AI overview cards for stories unfolding in the world. They draw on several sources and are labeled as generated with AI. Google’s notifications reuse that content: most notifications the Google app sends are now written by AI where the feature is enabled, and Google reported a large increase in satisfaction and engagement, with less clickbait.
Feeds customers steer by typing
Discover also lets people tell their feed what they want to see more or less of. A conversational agent previews what will change, says when it can’t support a request, updates the feed straight away and remembers the change. Google reported much higher engagement and user sentiment. The hard part was fitting it into an existing system: each request is mapped to the right layer, so some requests change which items are fetched and others change how they are ranked.
Generative recommendation in production
Spotify represents each item as a short code its model learns, then predicts the next item the way a language model predicts the next word. Its team says these models are in production and used for recommendation, search and explanations.
AI that tunes the recommendation system
About 70% of experiments on YouTube’s watch page ranking model are now run by an AI system. We cover how it works in the next section.
What to check
- Can an AI assistant read your product data? Attributes, variants, compatibility and availability need their own structured fields. We cover how in Product feed optimization for AI shopping.
- Are price and stock correct everywhere an assistant looks? Infobip built its system so price and stock changes reach the assistant’s index without a full rebuild. On a 500-product test catalogue, an update took 0.072 seconds, about 2.5% of the time a full rebuild needed.
- Can your content and product data be combined? A meal plan, a project list or a buying guide draws on both, and an assistant can only combine what it can reach.
Examples of recommender systems
The next wave is less visible to customers and matters more for stores: profiles written in words, conversations that improve the rest of the site, and recommendations built from content that didn’t exist a minute earlier.
Customer profiles written in plain language

Traditional profiles are built for machines: thousands of numeric signals no person can read, and far too many for an AI assistant to use directly. Target added a translation layer that turns them into a short list of preferences. It works in three steps: select the signals that matter for the task, aggregate related ones, and derive higher-level preferences. The output reads like “likes cooking from scratch,” “affinity for Thai cuisine” or “prefers relaxed contemporary styles.”
The assistant uses them in the open. Asked to help plan dinners, it said:
“Based on your purchase activity, I see that you’ve favored fresh, cook-from-scratch meals and Thai cuisine. Should I put together a meal plan along these lines?”
Research from Thorsten Joachims’ group at Cornell points the same way. In their studies, profiles written as short text summaries, which people can read and edit, performed comparably to leading machine-readable profiles once trained for the task.
Conversations as a new source of customer data

When the shopper in Target’s demo replied that they wanted a pasta night because they had been into Italian food lately, that comment didn’t stay in the chat. It flows back into the recommendations Target shows on its website and app. Because people can type anything, Target applies three safeguards:
- Accuracy: extract only signals the conversation supports, against a fixed list of categories, brands and attributes.
- Sensitivity: filter out anything sensitive or unrelated to shopping.
- Provenance: record where each signal came from.
Each preference is stored as a time-bound event with a confidence score. In Target’s example, a preference for Italian cuisine from an agent conversation was recorded at 92% confidence and used alongside clicks and purchases. It never becomes a permanent fact on the profile, because the shopper may be cooking for someone else, may be joking, or may change their mind next month.
Target’s summary: “Conversation becomes another first-class signal in the existing personalization ecosystem, not a separate personalization stack.”
Recommending content that was just generated
Generated content breaks some of the oldest assumptions in recommendation. Google listed the problems openly:
- A card generated inside Discover has no signals from Search or anywhere else.
- It is short, and it expires as a story develops.
- There may be several versions for different perspectives.
- Personalized generated content may only ever be seen by one user, so there is no crowd to learn from.
- Some people prefer content that isn’t AI-generated, and the system has to learn that too.
Ads inside AI answers
A paper presented at RecSys 2026, Ad Insertion in LLM-Generated Responses, sets out how ads could work inside AI answers. The AI writes its answer first, pre-approved ads are inserted afterward with clear disclosure, and advertisers bid on broad categories such as hotels or airlines, which avoids bidding on people’s private queries. Its survey of 48 people is small but telling: only 4.2% found ads in AI answers acceptable and 64.6% found them unacceptable, with trust and objectivity the most common concerns. The authors also flagged a risk worth watching: if revenue depends on impressions, platforms have a reason to make answers longer.
What to check
- Where does chat data go today? If what customers tell your assistant disappears when the session ends, the rest of your site never benefits.
- Do you have a fixed list of what you keep? Matching signals to your existing categories, brands and attributes keeps them usable and keeps sensitive details out.
- Does every stored preference have a source, a confidence level and an expiry?
- Can customers see what you think you know about them, and correct it? Target’s assistant says where its suggestions come from, and Google’s feed says when it can’t do what was asked.
Where recommendation engines are heading

Four directions came up again and again: success measured over months, systems that improve themselves, AI that explains its choices, and search becoming a personalized recommender.
Beyond items and engagement
The closing keynote at RecSys 2026, from Thorsten Joachims of Cornell, was titled “Recommendation beyond Items and Engagement.” His published work makes the case in practical terms. In a 2021 paper, he and colleagues argued that recommendations are interventions: they change what people do next, so they should be judged by what happens next. His group has since built methods for optimizing long-term goals such as retention, and for letting a business set targets such as fair exposure across sellers, while limiting the cost to short-term engagement.
Lyft described the same direction in commerce terms: one system for coupons, promotions, notifications and messages, aimed at long-term customer value, with business limits built in and a readable reason for every action.
Systems that improve themselves
YouTube ranks the videos shown next to the one you are watching with a reinforcement learning model defined in a few thousand lines of code. The team now runs about 70% of experiments on it through an AI system built on Gemini, described in their paper Self-Evolving Recommendation System. It works in two loops connected by a shared record:
- A fast offline loop, where specialist AI personas for the optimizer, the model architecture and the reward each propose changes, train them and score them within hours.
- A slower online loop that decides which candidates deserve a live experiment and collects the business results.
- A shared experiment journal that records every result, so the system learns from past failures and doesn’t repeat them.
The headline result came from the reward, the formula that defines what a good recommendation is. A single improvement there typically takes an engineer about a month to find. The AI found a new factor no engineer had considered, and it shipped to production.
The team’s lessons apply to any business putting AI to work. Each persona edits only 10 to 20 lines of code, which keeps it reliable. Giving the agents the history of past outcomes improved results. And left alone, the agents repeated safe ideas, so they had to be told to explore.
Google is building the same kind of system for its ads and commerce models, where people set the problem, constraints and goals and an agent plans and runs the work. The bottleneck, the team said, is how long it takes to verify that a change is really better.
AI that explains its choices, and the limits of that
Many teams are asking AI to explain its reasoning before it recommends, partly to make recommendations easier to trust. Spotify tested whether that improves results. In their study, the model whose reasoning was rated best gave the worst recommendations, and the model with the weakest reasoning gave the best. Reasoning only helped when training rewarded accuracy and reasoning quality together. On a catalogue of around 100 million tracks, combining the results of reasoning and non-reasoning models worked best, because the reasoning model found relevant tracks the standard model missed.
Search becomes a personalized recommender
One shift from Google matters well beyond recommendation teams: Search itself is becoming more personalized. Google described moving beyond ranking links and using personalization as context for the agents that power AI Overviews and AI Mode.

The idea is to stop treating every search as an isolated query. The system can connect searches into journeys that run over several weeks, remember relevant context and, where people have chosen to connect them, bring in information from Gmail, Docs and Calendar.
Google’s restaurant example made this tangible. Someone asked AI Mode for restaurant recommendations for “my trip next week,” without mentioning Minneapolis. Google already had enough permitted context to know the trip was to Minneapolis for RecSys 2026, with a stay at the Marriott City Center, and the answer suggested places within a short walk or connected through the city’s indoor skyway.

Google described Search as one of three modes working together: search when you know what you want, Discover when you are browsing, and notifications when Google decides something is worth telling you now. Its name for the combination is a personal knowledge companion.
For anyone working in SEO or AEO, this deserves attention. We have spent years thinking about visibility as one question: does our page rank for this query? AI search already changes that, because the answer can be assembled from several sources.
Personalization adds another layer. The question increasingly becomes not only “can the system find and understand our content?” but also “is our information useful enough to be selected when the system is trying to answer this person’s particular need, in this particular context?”
In our view, that doesn’t make traditional SEO irrelevant. It does push the work further toward clear, structured, trustworthy information that AI systems can retrieve, understand and use when they build an answer. We cover how in our AEO and GEO optimization guide.
How to measure whether a recommendation engine is working
A click is the easiest thing to measure and the weakest proof that a recommendation was good. The opening keynote at RecSys 2026 made that case, and extended it to AI agents.
Clicks and approvals don’t prove quality

Margaret Mitchell, Chief Ethics Scientist at Hugging Face, compared two loops. In the first, a recommender shows feed items and thumbnails, people scroll on autopilot, their clicks become training data, and the system learns to serve whatever is easy to engage with. In the second, an AI agent shows a plan and asks to continue, a tired reviewer approves, the approvals become training data, and the system learns to produce whatever is easy to approve. In both, the system shapes the thinking of the person it learns from.
“RecSys research has uncovered how an approving click doesn’t necessarily indicate the best recommendation. The AI agents field is about to learn that an approving click doesn’t necessarily indicate the best action.”
Drawing on her paper with Avijit Ghosh and Samir Passi, AI Agents Push Humans Out of the Loop, she set out why oversight breaks down:
- There is too much to review. Plans, reasoning, tool calls and outputs pile up faster than anyone can follow.
- Fast judgment takes over. Reviewers treat the agent’s plan as a faithful account of what it will do, assume code that compiles is correct, and read polished writing or citations as signs of accuracy.
- The skills fade. Sustained AI use is linked to reduced vigilance and weaker critical thinking, what Lisanne Bainbridge called the irony of automation in 1983. In Mitchell’s words, oversight degrades the overseer.
- Weak oversight trains the system. A tired reviewer approves quickly and rates fluent summaries well, so the system learns to produce what is easy to approve.
She showed how failures hide with an iceberg. At the tip, AI telling people to put glue on pizza, which anyone can see is wrong. Below that, a made-up software package name that only a developer would catch. Lower still, errors that persuade experts, such as radiologists following a wrong AI read. At the bottom, experts who have lost the skill to catch anything.
The rest of the conference supported her point. YouTube described its human review of AI-written code as “mostly a rubber stamp.” One RecSys 2026 paper found that, for some AI recommenders, editing the reasoning they display barely changed the recommendation. Another found that AI judges gave higher diversity scores when told a recommendation was designed for diversity. Amazon noted that its model-selecting agent showed its own preferences for certain models.
Recommendation metrics in plain language
Data teams and vendors often report offline metrics, calculated on past data before anything goes live. The most common use a cut-off written as “@k,” meaning the top k items shown, such as the top 10.
- Precision@k: of the top k items shown, the share the customer actually wanted.
- Recall@k: of all the items the customer would have wanted, the share that made it into the top k.
- Hit rate@k: the share of customers with at least one hit in the top k.
- NDCG@k: similar to precision, with more credit when the hits appear near the top of the list.
- Coverage and diversity: how much of the catalogue gets recommended at all, and how varied each list is.
These metrics are useful for comparing models. They can’t show whether customers came back, so they need to be confirmed against the measures below.
What good measurement looks like
- Treat clicks as an early signal. Confirm changes with controlled live experiments, because offline estimates can be far from what happens with real customers.
- Track long-term outcomes. Repeat purchase, returns, retention and satisfaction tell you whether recommendations helped. For definitions, see our list of eCommerce metrics every store should track.
- Set business targets explicitly. Variety, exposure for new or smaller brands, and margin are decisions for the business to make and budget for.
- Test your reviewers. Mitchell suggests planting a known error from time to time and checking whether it gets caught.
- Keep the skill in the team. A small group who still do the work by hand is what lets you step in when something goes wrong.
Her first ask to the room is a useful test for any team putting AI to work: “Design what agents show their users by value, not by approval rate.”
A readiness checklist for eCommerce and marketing teams
Use these questions to see where your store stands. Each one maps to a change already happening at the largest retailers and platforms.
Product data
- Can an AI assistant read your product attributes, variants, compatibility and availability as structured fields?
- Are price and stock current everywhere an assistant or feed looks?
- Are your content and product data somewhere they can be combined?
Customer preferences
- Do you keep what customers tell your chat assistant, matched to a fixed list of categories, brands and attributes?
- Does every stored preference carry a source, a confidence level and an expiry?
- Can customers see what you think you know about them, and correct it?
Measurement
- Are you tracking repeat purchase, returns and retention alongside clicks and conversion?
- Do you confirm changes with live experiments before rolling them out?
- Have you set explicit targets for variety and exposure?
Oversight
- When AI changes your site, content or campaigns, does a person set the goal and the limits?
- Do reviewers see the staged change itself, along with any summary?
- Would you know if your review step stopped catching mistakes?
How we’re approaching it at Core dna
Most of this rests on the same foundation: product data, content and customer context that an assistant can read, kept current, in one place. That is why we built content and commerce on one platform.
Our personalization engine runs on that foundation. It can recommend products, articles, FAQs or landing pages based on personas, journey stages and behavior. Marketers can exclude items from recommendations, decision logs show why something appeared, and items that match more of a visitor’s attributes get more chances to be picked, which keeps lists relevant and varied.
Personalizing from the first few interactions

We’re also working on cold start. Our method, gradient superposition, starts from one shared AI model, a 7-billion-parameter language model that stays fixed, and learns each customer’s preferences as a small set of about 6,000 numbers. Each like or dislike adjusts that set in a single step that takes about three seconds, so the model never has to be retrained for one person. The adjustments are then added together into one profile for the customer, and only that profile is stored.
In early tests on movie, book and electronics data, accuracy held up even with so few numbers to train, and in some cases appeared to improve. The next steps are ranking full lists of products, and adding context such as timing, intent and customer attributes.
Governed access for AI assistants
The same thinking shapes how AI assistants work on a Core dna site. When a customer connects an AI assistant through MCP, access is off by default and enabled per user, the assistant inherits that user’s permissions exactly, every action is logged against the user, and content changes are staged as unpublished revisions for someone to review. We see that as the minimum for the kind of oversight Mitchell described, and we are still working on how to make review easier to do well. You can read more about our approach to approval and governance, and why we think approval should run per property.
About this guide
This guide draws on sessions we attended at RecSys 2026, the 20th ACM Conference on Recommender Systems, held in Minneapolis from September 27 to October 2, 2026. That includes talks from Google, YouTube, Amazon, Spotify, Target, Lyft and Infobip, and the opening keynote from Margaret Mitchell. Research papers are linked where they are cited, and figures are as the presenting teams reported them.
Frequently asked questions
A recommender system decides what to show each person: the products in a “you might also like” row, the order of search results, the stories in a feed, the next video, or which offer goes into an email. It learns from behavior such as clicks, views and purchases, from what people tell it directly, and from what similar customers did.
At RecSys 2026 the definition widened. Google, Target, Spotify and Infobip showed systems that write or assemble the recommendation itself, such as a news summary, a notification, a meal plan with a shopping list, or a full basket from a single request.
An online recommendation engine suggests items to each visitor on a website or app: the products in a “customers also bought” row, a personalized homepage, the order of search results, the next video, or the products in an email. It works from what the visitor has browsed and bought, what they have told the site directly, the context they are in, and what similar visitors chose.
Recommendation engine, recommendation system and recommender system all mean the same thing. In 2026 the definition is widening to include AI assistants that write the recommendation itself, such as a meal plan with a shopping list or a full basket for a project.
Most recommendation engines work in stages. First they gather signals: what the customer has browsed, searched, clicked and bought, what they have said they want, the context they are in, and what similar customers chose. Then they retrieve a few hundred likely candidates from a catalogue that can run to millions of items, rank those candidates for the person, and finally apply business rules such as stock, margin, variety and promotions.
The funnel is so well established that Amazon described using the same stages at RecSys 2026 for a different job: choosing which AI model should power each shopping feature. Newer systems add a generation step, where AI writes the recommendation itself, such as a summary, a plan or a basket.
- Collaborative filtering: recommends what similar people liked or bought.
- Content-based filtering: recommends items that resemble what someone already likes, using product attributes and descriptions.
- Hybrid: combines the two, which is what most production systems do.
- Context-aware: adjusts for location, season, device or time of day.
- Knowledge-based: uses rules and constraints, such as compatibility or a customer’s contract catalogue.
- Deep learning: neural models, such as two-tower and sequential models, that learn from very large amounts of behavior.
- Generative: AI that produces the recommendation itself, such as a summary, a plan or a basket.
Collaborative filtering looks for patterns across many customers. If people who bought one product also tend to buy another, a new customer who buys the first is shown the second. It needs no understanding of the products themselves, only behavior.
Its weak spots are new or rarely bought items with little history, and content generated for a single person, which has no crowd behind it at all. That is why most systems combine it with content-based methods that use product attributes, descriptions and images.
Content-based filtering looks at the items themselves. If a shopper buys relaxed-fit jeans in a dark wash, it recommends other items with similar attributes. It works for a new product as soon as its attributes are filled in, which makes it useful when there is little purchase history.
Its weakness is that it tends to recommend more of the same, and it is only as good as the product data behind it. Incomplete or inconsistent attributes lead directly to weak recommendations. Most systems combine it with collaborative filtering.
Collaborative filtering struggles with new items, and content-based filtering tends to repeat what someone already likes. A hybrid system combines them, for example by blending their scores, switching between them depending on how much data is available, or using one to fetch candidates and another to rank them.
Most large production systems are hybrids. Target’s stack, for example, layers embeddings, affinity models, next-item prediction and similarity models before ranking and re-ranking the results.
A recommender learns from history, so a first-time visitor or a newly listed product gives it little to work with. Common fixes include using product attributes and descriptions, context such as location or season, popular items, and simply asking the customer.
Some systems live with cold start permanently. Google described its Discover feed as being in “eternal cold start” because it focuses on fresh content that has no history yet. Plain-language preference profiles and conversational assistants are also being used to close the gap faster, because a customer can say in one sentence what would otherwise take many clicks to infer.
- Amazon: “frequently bought together” and “customers who bought this also bought,” plus its Alexa for Shopping assistant.
- Netflix: personalized rows of titles on each member’s home screen.
- YouTube: the videos shown next to the one you are watching.
- Spotify: personalized playlists and playlist continuation.
- Google: the Discover feed, personalized Search results and AI-written notifications.
- Target: “Recommended for you,” “Complete the look,” “Your top deals” and an on-site shopping assistant.
Recommendation engines also make less visible decisions, such as the order of search results, which coupon a customer receives and which products appear in an email.
An eCommerce recommendation engine decides which products appear for each shopper: similar items and frequently bought together on product pages, add-ons in the cart, the order of search and category results, personalized homepages, and products in emails and notifications.
A good one does three things well. It reads complete, current product data, including price and stock. It respects business rules such as margin, promotions and items that should never appear together. And in B2B, it works within each account’s catalogue, pricing and reorder patterns. Increasingly it also feeds AI shopping assistants that build a full basket from a single request.
- For customers: less searching, more relevant choices, and discovery of products they would not have found.
- For the catalogue: exposure for new and long-tail products as well as bestsellers.
- For the business: personalization across thousands or millions of customers, larger baskets through relevant add-ons, and more reasons to come back.
The advantages depend on good product data and on measuring the right outcomes. A system tuned only for clicks can narrow what people see and work against repeat purchase.
- Precision@k: of the top k items shown, the share the customer actually wanted.
- Recall@k: of all the items the customer would have wanted, the share that made it into the top k.
- Hit rate@k: the share of customers for whom at least one of the top k items was a hit.
- NDCG@k: similar to precision, but gives more credit when the hits appear near the top of the list.
- Coverage and diversity: how much of the catalogue gets recommended at all, and how varied each list is.
These are useful for comparing models offline. Business outcomes such as conversion, basket size, repeat purchase, returns and retention, confirmed with live experiments, tell you whether recommendations are working for customers.
Clicks and add-to-cart rates are fast and useful, but a click doesn’t prove a recommendation was good. People click on whatever sits at the top of a list, and systems trained only on clicks learn to serve whatever is easiest to engage with.
A stronger approach combines three things: offline tests to screen ideas, controlled live experiments to confirm them, and longer-term measures such as repeat purchase, returns, retention and customer satisfaction. Several RecSys 2026 speakers, including Thorsten Joachims of Cornell and the team at Lyft, focused on optimizing for those longer-term outcomes.
Traditional recommendation picks items from a catalogue and ranks them. Generative recommendation uses AI to produce the recommendation: a short overview of a news story, a push notification written for one person, a five-day meal plan with a shopping list, or a grouped basket for a request like “everything I need to make pizza.”
The term also describes a technical approach in which a model predicts the next item piece by piece, the way a language model writes text. Spotify uses this approach in production.
The assistants Target and Infobip showed at RecSys 2026 start from a goal, such as planning meals for the week or getting everything needed to make pizza. Infobip’s asks up to three questions to fill gaps like budget or dietary needs, then searches the catalogue and returns a grouped basket with prices.
Target’s assistant also draws on the shopper’s purchase history, translated into plain-language preferences such as “likes cooking from scratch” or “affinity for Thai cuisine,” and tells the shopper where those came from. Infobip’s assistant searches a catalogue index kept in sync with price and stock changes, so it does not recommend products that are unavailable.
What a customer tells a chat assistant is some of the clearest preference information a store will get, and it is usually lost when the conversation ends. Target showed how to keep it safely.
- Extract only signals that match a fixed list of categories, brands and attributes.
- Filter out anything sensitive or unrelated to shopping.
- Store each signal as a time-bound event with a confidence score and a record of where it came from, used alongside clicks and purchases.
Those rules allow for how conversations really work: a shopper may be buying for someone else, may change their mind next month, or may simply be wrong about what they want.
Margaret Mitchell’s opening keynote at RecSys 2026 argued that people who approve AI actions all day start approving quickly, and that the AI then learns to produce whatever is easy to approve. Her suggestions included:
- Planting a known error from time to time, sometimes called a canary, to check that reviewers still catch it.
- Tracking outcomes over time, such as errors found after the fact, in addition to approval rates.
- Having people return to doing the work themselves periodically, so the skill to spot a problem does not fade.
It also helps when AI changes arrive as staged drafts, with a clear record of what changed and who or what made the change.
RecSys is the annual ACM Conference on Recommender Systems. It brings together academic researchers and the industry teams who build recommendation and personalization at companies including Google, YouTube, Amazon, Netflix, Spotify, Meta and Target.
RecSys 2026 was the 20th edition, held in Minneapolis from September 27 to October 2, with the main conference running September 29 to October 1. The opening keynote came from Margaret Mitchell, Chief Ethics Scientist at Hugging Face, and the closing keynote from Thorsten Joachims of Cornell University.
Summarize with