Skip to content
The Product Guys
All principles
Social

Reputation and Rating Systems

Ratings only discipline behaviour when raters are honest, and most designs quietly make them dishonest.


Reputation systems aggregate past behaviour into a signal that strangers can use to decide whether to transact. They work when ratings are informative, which requires raters to be candid, protected from retaliation, and representative of everyone who transacted. Most deployed systems fail at least one of those conditions and drift toward uniformly high scores.

How it shows up in software

Every marketplace, app store, gig platform, and creator tool ships some reputation surface. The design choices that matter are invisible to users: who may review, when reviews appear, whether the other side can retaliate, and how non-reviewers are counted. A five-star average built from a biased sample is worse than no average, because it looks like information.

Using it well

  • Reveal both sides simultaneously, or after a deadline, so nobody writes a review while fearing the reply.
  • Report the share of transactions that produced no review, since silence is the most informative part of the signal.
  • Verify that the reviewer actually transacted, and label verification in the interface.
  • Show a distribution rather than an average alone, so a split verdict does not hide behind a mean.

Where it turns manipulative

  • Buying or brokering reviews, and platforms that tolerate broker markets, defraud every buyer who reads the score.
  • Letting sellers filter or gate review requests, so only satisfied customers are invited to rate, is a rigged sample presented as a measurement.
  • Threatening or offering payment to remove a negative review coerces the signal; some platforms enable this by design when they allow retaliation.
  • Using rating thresholds to deactivate gig workers without appeal turns a noisy, biased number into an employment decision.

Where you have seen it

  • Airbnb

    Holds each party's review until both are submitted or the window closes, which is the simultaneous-reveal design tested in the Fradkin experiment.

  • eBay

    Publishes detailed seller ratings across several dimensions alongside the headline percentage, which breaks up a single compressed score.

  • Steam

    Shows both an overall and a recent review score and separates purchased copies from keys, so a burst of activity is visible rather than averaged away.

What the research says

  • Nosko and Tadelis, 2015Well evidenced

    Using eBay data, they showed that feedback left is a biased sample of experience, and that a measure built on the rate of silent bad experiences predicted future buyer behaviour better than the public score.

    Observational plus a controlled experiment at eBay. The core point, that unreported bad experiences carry the information, is robust.

  • Fradkin, Grewal and Holtz, 2021Well evidenced

    An Airbnb experiment that hid each party's review until both had submitted, or until a deadline passed, reduced retaliation and shifted the rating distribution downward, revealing previously suppressed negative reviews.

    Large randomised field experiment on the live marketplace.

  • Zervas, Proserpio and Byers, 2021Well evidenced

    Compared Airbnb and TripAdvisor listings for the same properties and found Airbnb ratings clustered at the top of the scale far more tightly, consistent with reciprocity and selection rather than with better accommodation.

Grades are a judgement about the evidence, not about how useful the idea is. Plenty of contested effects are still worth knowing, as long as you do not cite them as settled.

Related