{
  "id": 11240461,
  "title": "Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS",
  "url": "https://urgent.news/2026/10/01/uplifting-conversion-across-the-acquisition-funnel-with",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-01T16:51:04.000Z",
  "source": {
    "name": "AWS Machine Learning",
    "slug": "aws-machine-learning",
    "url": "https://aws.amazon.com/blogs/machine-learning/uplifting-conversion-across-the-acquisition-funnel-with-personalization-using-contextual-bandits-on-aws/"
  },
  "original_language": "en",
  "account": "Generative AI allows for quick creation of personalized content at low cost. However, the challenge arises in selecting the best content for each customer. Amazon Payments used a multi-objective contextual multi-armed bandit (MAB) on Amazon SageMaker AI to tackle this issue in an acquisition funnel. A seven-week A/B test showed a high single-digit percentage relative lift in final-funnel conversion for one customer group, while another group saw no improvement.\n\nMulti-armed bandits (MAB) are reinforcement learning methods designed for settings with many options and limited traffic. They learn which option performs best while continuing to serve customers. The key trade-off in MABs is between exploitation (serving the current best option) and exploration (trying less-certain options to gather evidence). UCB (Upper Confidence Bound) is a commonly used MAB strategy that balances both aspects.\n\nContextual MABs take this concept further by conditioning decisions on signals about the visitor. Instead of selecting the best overall content, contextual MABs consider features representing the visitor's attributes, such as payment behavior and transaction mix. This way, patterns learned for one context can be applied to similar visitors without the need for separate traffic in each segment.\n\nLinUCB (Linear Upper Confidence Bound) is a popular contextual bandit algorithm. It assumes that the expected reward for an arm is a linear function of the context vector, allowing it to generalize to visitors it hasn't seen before. LinUCB maintains two running tallies: one for rewarding visitors (b) and another for tracking the visitors the arm has seen (A). The estimate of the reward for each arm is calculated as θ = A⁻¹·b.\n\nTo optimize the entire customer journey with multiple stages, Amazon Payments ran separate LinUCB models for each stage (start, submit, and approve) and combined their UCB scores using linear combination weights. This approach optimizes the entire funnel simultaneously, addressing the seesaw problem where optimizing one stage can negatively impact others.",
  "summary": "Generative AI makes it cheap to produce personalized content at scale, but which variation do you show each customer? Amazon Payments used a multi-objective contextual bandit on Amazon SageMaker AI to personalize an acquisition funnel, achieving a high single-digit conversion lift for one audience, and learning why content, not the model, was the constraint.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}