A new report out this morning found that three barely-known websites built 215,128 machine-generated "best software" pages between them, and Perplexity's AI search engine treats those pages as trustworthy sources more often than it cites the software vendors themselves. Trellner Research published the findings a little after 8 AM UTC on September 2, after running 380 software-category questions through Perplexity's sonar and sonar-pro models and checking where the answers actually pointed.
- Trellner ran 380 software questions through Perplexity and pulled 7,534 citations from the answers.
- 59.8% of those citations point to domains ranked below the top 100,000 sites on the web, by Tranco's ranking.
- 23.4% cite domains with no meaningful ranking at all: sites Tranco doesn't register as having real traffic.
- Three sites, wifitalents.com, worldmetrics.org and gitnux.org, account for 215,128 of the pages behind those citations, all registered within five months of each other in 2023 and 2024.
What did Trellner actually measure?
The methodology is straightforward enough to check yourself. Trellner asked Perplexity's sonar and sonar-pro models, accessed through OpenRouter, 380 questions shaped like the ones a shopper actually types: "best project management software for small teams," "top CRM tools for real estate," that kind of thing. Then it collected every URL those answers cited as a source and ran each one against Tranco, the standard web-ranking dataset built from real browsing data, plus the Wayback Machine to check when each domain first appeared online.
RelatedClaude Code Adds a Built-In Browser That Reads the Live Web
The headline number is the median: across all 2,055 distinct domains cited, the typical one ranks 71,611th on the web. That's not a domain nobody reads, but it's nowhere near where you'd expect an AI engine to be pulling "best of" recommendations from. For comparison, a site most people would call a household name, in software review terms, sits in the low thousands or better. Perplexity's median source is two orders of magnitude further down the list.
Who is actually behind these sites?
Trellner names three: wifitalents.com, worldmetrics.org and gitnux.org, all registered between December 2023 and May 2024, all publishing the same kind of "Best X Services" roundup at industrial scale. Between them they account for the 215,128 generated pages behind Perplexity's citations. Checking the sites directly turns up a detail Trellner's report doesn't dwell on: gitnux.org has published a post titled "Is WifiTalents.com a Reliable Source?" and wifitalents.com has returned the favor with "Is Gitnux.org a Reliable Source?" Two sites under apparent common ownership vouching for each other's credibility is exactly the kind of signal a link-counting algorithm can be fooled by, and it's a step beyond just publishing volume.
A fourth domain shows up too, though it's a different case: guideflow.com, a real product's marketing blog, landed 194 citations, the third-highest of any single domain in the study. That's not manufactured content in the same sense, but it shows the same underlying weakness. Perplexity's citation logic doesn't appear to weigh whether a source is a neutral third party or a company blog with an obvious reason to be generous about the product it happens to be selling.
Why does Perplexity cite pages like this at all?
Sonar and sonar-pro work the way most retrieval-augmented answer engines do: the model doesn't "know" the best software for a given task off the top of its own training, so it runs a live search, pulls back a set of pages, and grounds its answer in whatever those pages say. That design is supposed to make answers more current and checkable than a model just guessing from memory. It only works, though, if the retrieval step actually favors pages with real editorial standing over pages built purely to be found.
Volume beats authority in that race more often than you'd hope. A site publishing tens of thousands of near-identical "best X" pages covers far more of the long tail of specific software questions than any single trade publication or vendor site could, so for a lot of niche queries, the manufactured page may genuinely be the only page that exists on that exact phrasing. Perplexity's model isn't choosing a bad source over a good one in that case; it's choosing the only source it found. That's arguably a worse problem than a ranking bug, because it means the fix isn't just re-weighting authority, it's making sure genuinely reliable sources exist to compete for that coverage in the first place.
RelatedGLM-5.3 tops CyberGym as Z.ai delays open weights two weeks
| wifitalents / worldmetrics / gitnux | Independent trade press | Vendor's own site | |
|---|---|---|---|
| Pages published | 215,128 combined | Dozens to low hundreds per year | One product page |
| Editorial review | None disclosed | Editors, disclosed methodology | Marketing-approved |
| Median Tranco rank | Far below 100,000 | Low thousands to tens of thousands | Varies by company |
| Financial interest in the answer | Unclear, possibly affiliate-driven | Disclosed if present | Direct, undisguised |
Who actually gets hurt by this?
Two groups, in different ways. Buyers researching software through Perplexity get recommendations shaped by whichever manufactured page happened to rank for their exact query, not by an editor who's actually used the product. And legitimate reviewers, from big trade outlets down to a single analyst who writes an honest comparison post, are competing for citation share against an opponent that can publish more pages in a week than they'll write in a decade. If AI answer engines become a meaningful share of how people find software, the sites that win that competition on volume alone get an outsized say in what gets bought, regardless of whether their advice is any good.
What should a reader do with an AI software recommendation right now?
Treat a Perplexity citation the way you'd treat an anonymous forum post, worth reading, not worth trusting blind. Click through to the actual source before acting on it, and specifically check whether the page reads like it was written by someone who used the software or like it was assembled from a template with the product names swapped in. Trellner's report gives a concrete tell: look at the page's own HTML title tag in the browser tab. A title like "Facts & Grounding Page" is not something a human writer produces; it's language aimed at whatever system is going to retrieve the page next.
- Whether Perplexity responds. The company hasn't commented on Trellner's findings as of publication; a citation-quality fix would be the meaningful signal, not a statement.
- Whether this generalizes past Perplexity. Trellner tested one engine because it exposes citations cleanly. ChatGPT search, Google's AI Overviews and Gemini all run the same retrieve-then-cite pattern, and there's no reason to assume they're immune.
- Whether more "shared operation" sites turn up. Once one cluster of co-registered, cross-vouching domains is documented, it's a template other operators can copy, or may already be running elsewhere unnoticed.
Our take
The uncomfortable part of this story isn't that someone gamed an algorithm, that's the oldest story on the internet. It's that an AI engine's citation is starting to carry the social weight of a fact-check, a little badge of "someone verified this," when Trellner's numbers show the underlying sourcing can be thinner than a random blog comment. Perplexity markets citations as the feature that makes it more trustworthy than a plain chatbot. A report showing the median citation sits 71,611 places down the web's authority ranking undercuts that pitch directly, and it's a problem every retrieval-based AI product shares to some degree, not just this one.
- ReportTrellner Research: Manufactured Sources Behind AI Recommendations
- Referencegitnux.org: "Is WifiTalents.com a Reliable Source?" example of the cross-vouching pattern between the named sites
Original analysis by GenZTech Team. Sources: Trellner Research.
