Skip to content
The Product Guys
All teardowns
YouTubeMetrics6 min read

What the watch-next rail is really optimising

A queue built from other people's sessions, not your intent.

The surface
The home feed, the sidebar or below-player rail of suggested videos, and autoplay on the watch page.
What the user wants
I came for one video and I want the next thing to be worth my time too.
01

Home before search

Opening YouTube lands on a recommendation grid rather than an empty search field, even for a viewer with a clear intent.

Supply-led discovery

Most sessions do not begin with a specific title in mind, so a search-first home would waste the majority of visits. Putting candidates in front of the viewer converts vague intent into a click. The cost is that viewers with a clear intent have to look past a wall of suggestions.

02

Thumbnail and title as a pair

Every candidate is represented by a creator-supplied thumbnail image and title, both chosen by the creator rather than generated by the platform.

Curiosity gap

A thumbnail that states the whole video has no reason to be clicked, so creators learn to withhold the resolution. The platform rewards whatever gets clicked and watched, so this shape emerges without anyone designing it. It is a market outcome, and it drifts toward overclaiming.

03

The watch-next rail

Suggestions beside or beneath the player are drawn heavily from what viewers with similar watch histories went on to watch, rather than only from the same channel or topic.

Collaborative filtering

Behaviour from similar sessions predicts the next click better than topical similarity, because topic misses mood, length and format. It also finds connections nobody would think to encode. The weakness is that it inherits whatever the crowd's attention was drawn to, including things nobody was glad to have watched.

04

Watch time as the target

YouTube has publicly described moving its recommender objective from clicks toward watch time and, later, toward satisfaction signals including surveys and explicit feedback.

Proxy metric drift

Clicks rewarded misleading thumbnails, so the objective moved to time watched, which rewarded padding and outrage. Every proxy gets gamed once it becomes the target, which is why the objective keeps having to move. The lesson is that the metric is a hypothesis about value, not value itself.

05

Autoplay after the video

When a video ends, the next suggestion begins after a short countdown unless autoplay is switched off.

Default as decision

The end of a video is the moment the viewer would otherwise leave, and autoplay removes that exit. Because the next item was chosen by the system, the viewer's continued watching is weak evidence that the recommendation was good. The signal and the intervention are entangled.

06

Not interested

Each suggestion carries a menu with controls to dismiss it or to stop recommending a channel, and YouTube surfaces history controls in settings.

Perceived control

A recommender the viewer cannot argue with feels like something happening to them. A dismissal control restores agency even when its effect on the model is small. It also captures the rare high-value negative signal, which passive behaviour never gives you.

The life of a proxy metric

1Pick a proxy you can measureWatch time, because satisfactioncannot be read off a server log.2Optimise it hardIt works. The proxy and the real thingstill move together.3Somebody learns to move the proxyaloneLonger videos, cliffhangers, titlesthat promise a payoff at minute nine.4Patch the objectiveSurveys, dislikes, not interested. Asecond signal that the first onecannot fake.
Watch time was a reasonable stand-in for whether people valued a video, and it stayed reasonable right up until it became worth money to move. The pattern is not unique to recommendations: plan the next version of the objective before the current one starts to rot.

What not to copy

  • Watch time as an optimisation target rewards length and emotional pull over usefulness, and no amount of downstream survey patching fully removes that pressure.
  • Autoplay contaminates the very signal the recommender learns from, because continued watching after a system-chosen video is not the same as choosing it.
  • Negative feedback is buried in a per-item menu while positive feedback is a single tap, so the model hears approval far more readily than objection.
  • Creators are optimising against the same ranking the viewer experiences, which means thumbnail overclaiming is a structural outcome and not a moderation problem.

The takeaway

Every proxy metric you optimise will be gamed by someone, so plan the objective's next version before the current one starts to rot.

Finished the teardown? Bank it and the day counts toward your run.

Where the principles come from

Written from public behaviour of the product, not from inside it. Interfaces change often, so treat the flow described here as of the time of writing and check the live product before quoting it.