Knowledge

How to Track Brand Mentions Across Social, Forums & Review Sites

Mention tracking is a coverage problem before it is a sentiment problem. How to build query sets, dedupe syndication and see mentions outside your own region.

Matt Brown

Matt Brown

September 6, 2026 · 8 min read

Every brand monitoring tool produces a chart of mentions over time, and the chart is only as honest as the coverage behind it. A dip can mean people stopped talking about you. It can equally mean a forum started rate-limiting your collector, or a review site began serving a different catalogue to the region you happen to collect from.

The two look identical from inside the dashboard. That is the central problem in mention tracking, and it is a coverage problem before it is a sentiment problem.

The surfaces behave differently

Treating “the internet” as one source is what makes most in-house monitoring fragile. Each surface has its own access model, its own structure and its own failure mode.

Social platforms. Where official APIs exist and cover what you need, use them. They are stable, they are permitted, and they will not silently change your coverage. Their limits are usually historical depth, rate caps and which fields they expose.

Forums and communities. Reddit-style aggregators, niche technical boards, industry-specific communities. Structurally simple, often the highest-value mentions, and the most fragmented. There is no single access route, so this is where per-source work accumulates.

Review sites. Software directories, retail reviews, app-store reviews. Heavily regionalized, which is the trap covered below. Review content is also structured, so it is the easiest surface to score and the easiest to over-weight.

News and blogs. Well served by existing feeds and aggregators. Usually not worth building bespoke collection for.

Question and answer sites. Low volume, high intent, and worth watching precisely because a question about your product often precedes a decision.

The practical advice is to rank surfaces by where your audience actually is rather than by how easy each one is to collect. A monitoring program that covers the two easy surfaces well and the important one not at all is a common and expensive outcome.

Geography changes what a mention even is

This is the failure that catches teams with international customers.

App stores serve different catalogues and different review sets by storefront. Software directories localize both listings and reviews. Social platforms vary what is visible by region. A mention posted by a customer in Germany can be genuinely invisible from a US collector, and nothing about the result will indicate that something is missing.

The consequence is that a single-vantage-point collector produces a regional view labelled as a global one. If you report “mentions are up 12 percent” to a leadership team with international revenue, that number should have been collected from the markets that revenue comes from.

Residential proxies with country and city targeting are what make a regional view actually regional. With the Shifter gateway, the vantage point and session go in the credentials against p.shifter.io:443:

customer-USERNAME-country-de-sid-mentions-de-ttl-600:PASSWORD

Hold one session per market and query so that a paginated thread or review list stays internally consistent, rather than rotating mid-collection and stitching together pages from different vantage points. The product view of this workflow is on the brand monitoring proxies page.

Build the query set deliberately

Coverage is bounded by what you search for, and most teams search for too little and then for too much.

Start with the obvious: brand name, product names, domain. Then add what real people actually type: common misspellings, the name without spaces, the abbreviation, the legacy name if you rebranded. Then add the contexts where you appear without being named directly: comparisons against competitors, category terms plus a complaint, executive names.

The counter-problem arrives if your brand name is a common word. A generic name buries real mentions under noise, and the fix is not a better model, it is a tighter query: require a co-occurring term, restrict to relevant communities, or match the domain rather than the word. Decide this at query design, because a noisy corpus poisons everything downstream.

Write the query set down as a versioned artifact. When your mention volume steps up next quarter, you need to know whether the world changed or your queries did.

Deduplicate before you count

The same statement reaches you many times. A press release syndicates across a dozen outlets. A social post is quoted, screenshotted and re-shared. A review is mirrored by an aggregator. A forum thread is cross-posted.

Counting all of those as separate mentions inflates volume in a way that correlates with syndication rather than with attention. Worse, it makes a single loud event look like broad sentiment.

A workable rule is a composite key of normalized author, normalized text and a time window, with a canonical-source preference so the original outranks the copies. What matters more than the specific algorithm is that it is fixed and documented, because changing dedupe logic retroactively rewrites your own history.

Keep the copies linked to the canonical record rather than discarding them. Spread is a real signal, it is just a different one from volume.

Cadence should follow risk, not uniformity

Collecting everything hourly is expensive and mostly wasteful. Collecting everything weekly means finding out about a problem after it has resolved itself badly.

Tier the surfaces. The places where a complaint escalates fast, typically social and the most active community for your category, justify a short interval. Review sites and directories move slowly enough for daily. News aggregation is usually near-real-time already through feeds.

Then add an event-driven layer: a launch, an outage, a pricing change or a press cycle should temporarily raise the cadence on the surfaces where that story would spread. A fixed schedule cannot do this, and the moments it misses are exactly the ones the program exists for.

Measure your own coverage

The discipline that separates a monitoring program you can trust from a chart you cannot: track the collection alongside the mentions.

Record success rate per source, results returned against results expected, and pages collected against pages available. When mention volume drops, the first question is whether your own success rate dropped with it. If it did, that is a collection story, and reporting it as a decline in conversation would be wrong.

Run a small control query against each surface twice per collection window from two different exits in the same region. Convergence means your view is stable; divergence means the source is responding to something about your requests. The method for establishing that baseline is in testing proxy speed, success rate and location accuracy, and the request-pacing side is in rate limiting and request throttling.

What to do with what you find

Monitoring that produces no routing decision is a report nobody reads. Three categories are enough to start.

Act now. A specific complaint from an identifiable customer, a factual error about your product, a security claim. Route to a human with the source link.

Aggregate. Sentiment, volume, share of voice against competitors. Weekly, in trend form, with coverage metrics attached so the trend can be believed.

Archive. Everything else. Searchable, not surfaced.

The reputation-defence side of this, including where mention tracking meets counterfeit and misuse detection, is covered in using proxies to protect your brand. App-store specifics are in scraping App Store and Google Play data.

FAQ

Should I use APIs or collect public pages?

APIs first, wherever one exists and covers what you need. Public-page collection fills the gaps, which in practice is most forums, most review sites and most long-tail communities.

How far back should historical collection go?

Far enough to establish a baseline you can compare against, which usually means a quarter. Backfilling years is expensive and rarely changes a decision.

Why does my mention volume jump when nothing happened?

Usually syndication of a single item, or a query change. Check the dedupe result and the query version before concluding anything about attention.

Do I need this if I already have a paid monitoring tool?

Ask the tool what its coverage is per surface and per region. Most in-house work exists because a specific important community, or a specific market, is not covered by the tool.

The bottom line

Mention tracking fails quietly. Coverage gaps look like silence, syndication looks like consensus, and a domestic vantage point looks like a global view.

Build the query set deliberately, dedupe on a rule you can defend, tier cadence by risk, collect from the markets your customers are actually in, and measure your own coverage alongside the mentions themselves. That is what makes the chart worth showing to anyone. Plans and rates are on the pricing page.

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started