Fake reviews are no longer just a reputation problem. They are now a legal one in the United States, the United Kingdom and the European Union, and the platforms that host reviews remove them by the hundreds of millions. For a brand, a marketplace or a trust and safety team, that raises a practical question: how do you find review fraud across thousands of products, not just the one listing someone complained about?
This guide covers the scale of the problem, the signals that actually separate campaigns from customers, tested code for the three most useful ones, and how to collect review data responsibly.
Key takeaways
- The scale is large. Google blocked or removed over 292 million policy-violating reviews on Maps in 2025, and Amazon says it stopped more than 250 million suspected fake reviews in 2023.
- Reading reviews does not work. In a well-known study, human judges spotting fake hotel reviews performed close to chance, between 53% and 62% accurate, while a trained classifier reached about 90%.
- Coordinated campaigns leave behavioural traces: bursts of reviews in a few days, near-identical wording, and reviewers with no other history.
- Combine signals. In our test, wording similarity alone was noisy, while bursts combined with duplicate wording or one-review accounts caught every planted fake review with at most three false positives.
- Treat flags as leads for human review, pseudonymise reviewer data, and check the legal position before acting on findings.
How big the problem is
The platforms’ own transparency figures give a sense of scale:
| Platform | Figure | Period |
|---|---|---|
| Google Maps | Over 292 million policy-violating reviews blocked or removed; over 13 million fake Business Profiles removed | 2025 |
| Amazon | More than 250 million suspected fake reviews stopped | 2023 |
| Tripadvisor | 2.7 million fraudulent reviews out of 31.1 million submitted; 54% of the fraud was businesses boosting themselves; 214,000 AI-generated reviews removed | 2024 |
Those are the reviews that were caught. They also show what fraud mostly looks like: on Tripadvisor, more than half was review boosting by businesses or people connected to them, not elaborate deception.
The rules have changed
Three major markets now prohibit fake reviews explicitly:
- United States. The Federal Trade Commission’s rule on consumer reviews and testimonials, in force since October 2024, prohibits fake reviews, buying positive or negative reviews, undisclosed insider reviews, company-controlled review sites presented as independent, and review suppression, and allows the agency to seek civil penalties.
- United Kingdom. Since 6 April 2025, submitting or commissioning fake reviews, concealing incentivised reviews and publishing reviews in a misleading way are banned practices, and businesses that publish reviews must take reasonable steps to prevent and remove fake ones. The regulator can fine up to 10% of global turnover.
- European Union. Since May 2022, traders that show reviews must say whether and how they check that reviews come from customers who actually used or bought the product, and submitting or commissioning fake reviews is prohibited.
For marketplaces and review publishers, that turns detection from a nice-to-have into part of compliance. For brands, it creates a route to act when competitors buy reviews. This is general information, not legal advice; consult counsel on your own situation.
Why reading the reviews does not work
The best-known academic study of the problem, by Ott, Choi, Cardie and Hancock at Cornell, built a set of 800 hotel reviews: 400 genuine and 400 written to deceive. Three human judges classified them. Their accuracy was 53.1%, 56.9% and 61.9%, and for two of the three the difference from guessing was not statistically significant. All three leaned towards believing reviews were genuine. A classifier trained on word patterns reached 89.8%.
That has two implications. Manual spot checks will miss most fakes. And text-only detection is now harder than it was in 2011: a model can write varied, natural-sounding reviews on demand, which is why Tripadvisor now reports AI-generated reviews as their own category. The signals that hold up better are behavioural: when reviews arrive, who writes them, and how similar they are to each other.
The signals that hold up
| Signal | What it catches | Weakness |
|---|---|---|
| Bursts | Dozens of reviews in a few days on a product that normally gets two a day | Genuine spikes after a sale, a launch or press coverage |
| Near-duplicate wording | Templated reviews with small word swaps | Common phrases in genuine reviews; varied AI-written text |
| One-review accounts | Reviewers created for a single campaign | New genuine customers |
| Rating shape | A sudden run of five-star reviews with little detail | Products that really are good |
| Cross-product overlap | The same reviewers appearing on one seller’s products | Loyal customers |
| Cross-market differences | A product rated very differently in one country’s storefront | Real differences in local versions or service |
No single signal is reliable. A burst can be a successful launch; similar wording can be a common phrase. Campaigns tend to show several signals at once, which is what makes combining them effective.
The code
The module below implements three of these signals and one way to combine them. It needs no external libraries. Reviewer names are replaced with salted hashes before analysis, since reviews are personal data and the analysis never needs the name itself.
import hashlib
import re
from collections import Counter, defaultdict
from statistics import median
def pseudonymise(author, salt):
"""Replace a reviewer's name with a stable token, so analysis never needs the name itself."""
return hashlib.sha256((salt + author).encode()).hexdigest()[:12]
def review_bursts(reviews, k=6.0, min_count=5):
"""Flag days whose review count is far above the product's typical day, using a robust baseline."""
daily = Counter(r["date"] for r in reviews)
counts = list(daily.values())
base = median(counts)
spread = median(abs(c - base) for c in counts) or 1
flagged = []
for day, n in sorted(daily.items()):
if n >= min_count and n > base + k * spread:
day_reviews = [r for r in reviews if r["date"] == day]
five_star = sum(r["rating"] == 5 for r in day_reviews) / n
flagged.append({"date": day, "reviews": n, "baseline": base, "five_star_share": round(five_star, 2)})
return flagged
def shingles(text, size=5):
text = re.sub(r"\s+", " ", text.lower()).strip()
return {text[i:i + size] for i in range(max(1, len(text) - size + 1))}
def near_duplicates(reviews, threshold=0.5):
"""Pairs of reviews by different reviewers whose wording overlaps heavily (Jaccard on 5-character shingles)."""
sets = [shingles(r["text"]) for r in reviews]
pairs = []
for i in range(len(reviews)):
for j in range(i + 1, len(reviews)):
if reviews[i]["reviewer"] == reviews[j]["reviewer"]:
continue
overlap = len(sets[i] & sets[j]) / len(sets[i] | sets[j])
if overlap >= threshold:
pairs.append((reviews[i]["id"], reviews[j]["id"], round(overlap, 2)))
return pairs
def one_review_accounts(reviews, history):
"""Share of reviewers in this set who have reviewed nothing else; history maps reviewer to their total reviews."""
reviewers = {r["reviewer"] for r in reviews}
return sum(history.get(x, 1) == 1 for x in reviewers) / len(reviewers)
def suspicious_reviews(reviews, history, threshold=0.5):
"""Combine the signals: a review is suspicious when it sits in a burst and either duplicates
other wording or comes from a one-review account."""
burst_days = {b["date"] for b in review_bursts(reviews)}
duplicated = defaultdict(int)
for a, b, _ in near_duplicates(reviews, threshold):
duplicated[a] += 1
duplicated[b] += 1
return [r["id"] for r in reviews
if r["date"] in burst_days and (duplicated[r["id"]] or history.get(r["reviewer"], 1) == 1)]
The burst detector uses the median and the median absolute deviation rather than the mean and standard deviation, so the burst itself does not inflate the baseline it is measured against. The pairwise comparison in near_duplicates is fine for a single product’s reviews; across a whole catalogue, group reviews by product or seller first, or use locality-sensitive hashing.
How it performed
We tested the code on a constructed dataset where the answer was known: one product with six months of organic reviews, between 256 and 285 of them at a few a day with realistic ratings and varied wording, plus a planted campaign of 40 templated five-star reviews from new accounts over three days. We ran it with six different random seeds.
| Check | Result across six runs |
|---|---|
| Campaign days found by burst detection | All three, every time; 13 to 16 reviews a day against a baseline of 2 |
| Planted reviews caught by the combined rule | 40 of 40, every time |
| Genuine reviews wrongly flagged by the combined rule | 0 to 3 |
| Near-duplicate pairs involving a genuine review, wording alone | 27 to 42 |
| Reviewers with no other reviews | 81% to 100% on campaign days, against 27% to 31% overall |
The last two rows are the important ones. Wording similarity on its own flagged dozens of pairs involving genuine customers, because real reviews share ordinary phrases. Only when it was combined with a burst did it become precise. That is the pattern to copy: let one signal nominate and another confirm.
A constructed test is not the real world. Real campaigns spread reviews over weeks, vary the text, and age their accounts before using them. Expect lower recall in practice, tune the thresholds on cases you have already confirmed, and treat every flag as a lead for a human reviewer, not a verdict.
Collecting review data
Detection needs the reviews, with dates, ratings and enough reviewer context to see history. A few points matter when collecting them:
- Collect across markets. The same product often shows different reviews in different country storefronts, and campaigns are frequently aimed at one market. Collecting each storefront from inside its country shows what local shoppers see, the same approach used for monitoring customer reviews.
- Keep history. Bursts and one-review accounts can only be measured over time, so collect on a schedule and keep what you collected.
- Minimise personal data. Store pseudonymised reviewer tokens rather than names where you can, keep only the fields the analysis uses, and set a retention period, as covered in residential proxies and GDPR.
- Stay within the rules. Collect public review pages at a modest pace, respect each site’s terms, and leave anything behind a login alone.
Review fraud rarely travels alone. Sellers who buy reviews often also hijack listings or sell counterfeits, and review text increasingly needs the same scrutiny as any other AI-generated content. For teams running this across many sources, the wider workflow is described on our content moderation at scale page.
The bottom line
Fake reviews are common, now explicitly illegal in the largest markets, and hard to spot by reading. What gives campaigns away is behaviour: too many reviews too quickly, too similar to each other, from accounts with no other history.
Collect reviews over time and across markets, pseudonymise who wrote them, combine signals so each confirms the others, and send what you flag to a person before anyone acts on it.
Sources and references
- Google, New ways we’re protecting businesses on Maps, 2025 figures.
- Amazon, How Amazon maintains a trusted review experience.
- Tripadvisor, 2025 Transparency Report.
- Federal Trade Commission, final rule banning fake reviews and testimonials, August 2024.
- Ott, Choi, Cardie and Hancock, Finding Deceptive Opinion Spam by Any Stretch of the Imagination, ACL 2011.
- Constructed test dataset and code by Shifter, 2 October 2026.