Last week we ran a real micro-SaaS through 5 AI assistants (GPT, Gemini, Perplexity, GLM, Qwen) and asked questions a potential customer would ask: "best AI launch coach for first-time founders", "top tools in 2026", etc.
The tool is genuinely good and it shows: 4 of 5 AIs recommended it, mostly at position #1.
Then the cracks:
Qwen (China) has zero knowledge of this product. Not one mention across 5 questions. A whole AI ecosystem — used by Chinese developers and creators — simply doesn't know it exists.
In a head-to-head ("Zarek vs LaunchList — which is better?"), Perplexity picked the competitor. The stated reason: "no verifiable public information about a product called Zarek."
Here's the kicker: when Perplexity did cite the product, the only source was the product's own site. AI models default to "trustworthy = independently verifiable." Your own homepage is the weakest citation you can have.
Three things any solo founder can do about this:
Give the models a second, independent source. A Dev.to post, a Product Hunt listing, a directory entry — anything not on your own domain. This was the single highest-leverage fix in the report.
Don't ignore the Chinese AI ecosystem. Qwen/GLM/Kimi are trained mostly on Chinese sources. One Chinese-language post or directory listing closes that gap entirely.
Say your category the same way everywhere. The product won on "AI launch coach" searches but lost "waitlist tool" searches because that's the competitor's category. Consistency in how you describe yourself is a ranking signal for AI.
We built BrandScope to run exactly this kind of check — it asks 6 AIs live questions about your brand and shows what they actually cite (not just whether they mention you). The free check takes 30 seconds:
brandscope.dev
Question for you: have you ever checked what AI assistants say when someone asks them to recommend tools in your category? It's usually not what you assume.
The Perplexity line is the bit that stuck with me: your own homepage is the weakest citation. I'd been treating directory listings as busywork. If the models only trust a second source, then one decent independent writeup is doing more than polishing the landing page again.
This gives me a concrete way to test a rollout I'm doing now. We have a consumer AI home-design product and are deliberately creating two kinds of independent sources: short directory profiles and deeper English/Japanese articles. My working hypothesis is that a directory fixes the existence layer — “this product is real and citable” — while a detailed article supplies the use-case language that changes how and when the model recommends it. I'm going to rerun the same prompts after each source is indexed instead of adding ten links at once, otherwise attribution becomes impossible. One variable I'd love to see in your data: did you control for indexing age? A source published yesterday and a source crawled for months can look like a quality difference when it may actually be retrieval delay.
You've made a stronger case for third-party evidence than another GEO checklist. I keep seeing founders, myself included, treat homepage copy like it has the same credibility as a neutral source. Building DictaFlow taught me that product proof needs to live somewhere outside your own site. People need to see the product working in a real workflow, not just read my claim. But how much is a third-party mention worth when it's clearly just a thin directory page?
The "works in a real workflow" point is the sharpest thing in this thread — that's a different category of evidence than "has a landing page," and your DictaFlow experience is the honest version of it. But our data says thin directories still pay off, at a different layer:
In the run behind this post, GPT recommended Zarek 5/5, and 4 of those answers cited thin directory pages (hunted.space, debutly.app profiles) — not case studies, not walkthroughs. A thin page is enough to flip the model from "no verifiable information" (the exact reason Perplexity initially picked the competitor) to "independently citable." It buys the existence layer: whether the model dares to recommend you at all.
What it doesn't buy is the persuasion layer — how the model talks about you. That's where real-workflow proof matters: a writeup or demo gives the model narrative and specifics it can echo, while a directory entry gives it nothing to say beyond "it exists." Most recommendation losses we see aren't "never heard of me" — they're "heard of me, but nothing meaningful to cite."
Honest caveat: we tag every citation own-site vs third-party per mention, but we don't yet grade third-party sources thin vs substantive. Your question is basically the spec for the layer we want to build next.
Do you think the models reward substantive sources at all — or is "exists + citable" all they're pattern-matching on? Our data leans toward the latter so far.
The second-source advice matches what we've seen — but there's a measurement gap: once you plant those Dev.to or PH citations, most analytics can't tell you which AI actually cited you or what that visit does next. We built https://amami.dev to split AI-referrer sources so you can see if the GEO work pays off.
The measurement gap is real, and I think the two tools sit on different halves of it. Amami splits AI-referrer traffic — the click side: which engine actually sent a visitor, and what they did after. We measure the answer side: what each engine says about a brand, and what it cites — no click required. The uncomfortable truth is that most GEO budgets get judged on one side while the loop spans both: you plant third-party sources (answer side), a user clicks through an AI answer (click side), and analytics loses the thread in between.
There's a fun data point hiding at the intersection if you ever test it: engines that cite you but never send traffic, versus engines that send traffic without citing you. Every brand we've run has both columns, and they're never the same engines.
Curious what the click side looks like in your data across GPT vs Perplexity — does the engine that cites best also send the most?
This matches almost exactly what I see doing GEO audits for B2B SaaS clients, the third-party citation point especially. Your own homepage being the worst possible source makes sense once you think about it from the model's side: it has no way to verify a company's own claims about itself, so it treats self-published pages as marketing, not evidence. The fix is usually cheap (one Product Hunt listing, one Dev.to writeup, one directory entry) but almost nobody does it because it doesn't look like "real" SEO work.
The category-consistency point is the one people underestimate most, though. I've seen the same thing: a product ranks well for the phrasing it uses about itself and just doesn't exist for the phrasing its buyers actually use, because the model is pattern-matching against a category cluster, not doing semantic search. That's usually a bigger swing factor than backlinks or citation count.
The Qwen/GLM blind spot is a good catch too — most Western founders never check anything past ChatGPT/Gemini/Perplexity, and if you sell into any market with a meaningful Chinese dev/creator audience, that's a real gap, not a nice-to-have. Did you notice whether Kimi behaved differently from Qwen/GLM, or did all three Chinese-trained models miss the product for the same reason (no Chinese-language source at all) ?
This is the most useful read of the post so far — "self-published pages as marketing, not evidence" is sharper than our own phrasing. And you're right that the cheap fixes look like nothing, which is exactly why almost nobody does them: a directory listing reads as busywork, not SEO, so it never makes the roadmap.
On category consistency — fully agree it's the one people underestimate. In the underlying data, Zarek wins the phrasing it describes itself with ("AI launch coach", ranked #1 across most engines) and is invisible in the phrasing its buyers actually use ("waitlist tool" — the competitor's category). The model is matching a category cluster, not searching semantically, and it's the cheapest thing to fix: a copy change, not a link change.
On Kimi vs Qwen/GLM — they don't behave uniformly, at least in our runs. Latest full run (6 engines × 5 questions): GLM and Qwen went completely blank on all five. Kimi didn't — it mentioned Zarek once, but the context was "no verifiable information about Zarek": it knows the name exists, then rejects it for lack of data. And Qwen is stranger still: it mentioned Zarek repeatedly in earlier runs (Aug 19 and Aug 21), then went totally blank. So it's not one single "no Chinese-language source" failure: GLM is a clean zero, Qwen forgets what it knew, Kimi half-knows. That pattern suggests retrieval here is snapshot-driven and unstable, not just a corpus gap.
Curious — in your client audits, have you seen Chinese-trained models split this way (one knows, one forgets, one half-knows), or is that just our small sample?
The point about homepage copy being the weakest citation for LLMs hit hard. I spent way too much time tweaking my landing page text thinking ChatGPT or Perplexity would index it, but realized they basically ignore self-claims unless there's third-party proof on places like Dev.to or Product Hunt. Definitely going to put up a few external directory listings this week.
That's exactly the pattern our data shows — and you're ahead of most founders on it: the usual version is spending months on the landing page copy, then wondering why no model ever picks it up.
One refinement before you spend the week on directories: it's per-model. In the run behind this post, GPT cited third-party listings (hunted.space, debutly.app) for 4 of 5 answers — the directory path is real, at least for GPT. Perplexity, on the exact same footprint, anchored every answer on zarek.tech alone and never touched the third-party pages. So "get listed" is right, but which directory matters less than whether that model actually retrieves from it — a nofollow'd corner of the web may as well not exist for some engines.
Practical version: list where the models you care about demonstrably look. And after you post them, a 30-second check will show which engines actually picked it up — that's the whole problem we built this for.
What are you planning first — a Dev.to writeup, or directories like Toolify / TAAFT?
Did the four AIs that did recommend it cite anything besides the product's own site, or was that only Perplexity's complaint? And was Qwen's blank a freshness gap or a coverage gap — would a listing on a big Chinese aggregator have fixed it, or is the model just months behind regardless of what you publish? Curious too whether you re-ran the same five questions later, and if BrandScope shows which citation eventually displaced the homepage rather than just whether the mention appeared.
Thanks — three good ones, and we can answer all three from actual data.
Citations: it wasn't homepage-only, and the "complaint" is model-specific. In the run behind this post (6 AIs × 5 questions, latest full run): Perplexity recommended Zarek 5/5 but anchored every answer on zarek.tech — that one is a self-citation story. GPT also recommended it 5/5, and four of five answers cited third-party listings (hunted.space and debutly.app both carry a Zarek profile — OpenAI clearly retrieved those, not us). Gemini cited via Google grounding links, which mask the final domain on our end. That's exactly why we tag every source as own-site vs third-party per mention instead of averaging it away.
Qwen: it's a coverage/stability gap, not a freshness gap. Qwen has mentioned Zarek in earlier runs (Aug 19 and Aug 21), then went completely blank on the latest run. So it's not "the model is months behind no matter what you publish" — it knew the product and then didn't. Would a Chinese aggregator listing fix it? GPT is the proof that syndication works (it found Zarek via directory pages, not the site). Qwen's retrieval of English indie products is snapshot-driven and spotty; a Chinese-language anchor (36Kr / Juejin / Zhihu write-up, for example) would plausibly raise its recall. Honest caveat: we haven't A/B'd it, so that's "probably yes", not "proven yes".
Re-runs and citation evolution: We ran the same 5 questions 6 times between Aug 19 and Aug 24 (three times on Aug 21 alone — debugging). Nothing since then, so no — we haven't re-run after publishing. And "which citation eventually displaced the homepage" — we don't show that yet. Every run stores the full source URL per mention, so the data for a citation-diff view already exists; the view itself (which source displaced which, per question) is the next layer we haven't shipped. Right now it's "did they mention you, and what did they cite" — not "what replaced what."
The Perplexity/GPT split is the most interesting number in there, and I think it slightly undercuts the "give the models a second source" framing from the original post. Zarek's third-party footprint was identical for both models — same hunted.space and debutly.app profiles sitting there — and GPT found them 4/5 while Perplexity never left zarek.tech. That makes the self-citation a retrieval property of the model, not a property of your footprint. The actionable version isn't "get a second source," it's "get a source inside the corpus this particular model retrieves from," which is a much harder instruction to give a founder — and probably worth separating in the report rather than collapsing into one recommendation.
On Qwen: if it's Aug 19 yes → Aug 21 yes → latest blank, then those three same-day runs on Aug 21 are accidentally your most valuable data. If they disagreed with each other, you have a within-day variance figure, and that bounds how much any single run can tell anyone. It would also argue the headline metric should be a mention rate over N runs rather than a binary mentioned/not — which sounds cheaper to ship than the citation-diff view, and it makes the "went blank" result interpretable instead of alarming.
One note on the Chinese anchors, since that's the side I work on: 36Kr / Juejin / Zhihu are not interchangeable in cost. 36Kr realistically needs PR or paid placement; Juejin and Zhihu are self-publish, so those two are the cheap A/B you're missing rather than a project. Worth knowing going in that both wrap or nofollow outbound links pretty aggressively — fine if you only care about model recall, useless if you were hoping the same post did SEO work at the same time.
You're right, and the sharper version of this is that the recommendation layer of the post is over-generalized even though the data layer isn't.
Our report already tags every citation as own-site vs third-party per model — GPT found Zarek via hunted.space and debutly.app; Perplexity never left zarek.tech. Same footprint, different retrieval. So the data was always model-specific. The post then averaged six models back into one instruction ("add a second source"), and that's the part your critique correctly takes apart: a source only helps if it's inside the corpus that model actually retrieves from. The honest instruction is per-model — "get a source into the corpus Perplexity (or whoever) pulls from" — which is a much harder ask for a founder, and probably deserves its own card in the report instead of one generic bullet.
On Qwen and the same-day runs: agreed, and the three Aug 21 runs in our logs don't even agree with each other — Qwen shows true in some, blank in others. That's the within-day variance you're describing, and it's exactly why our weekly report carries a note on every card: single snapshot, ±10% is normal, trend beats any single week. A blank run alone is almost never a signal about the brand; it's only interpretable against multiple runs. Making the headline a rate over N runs is the obvious direction — the only reason it isn't shipped yet is cost: every model × question × N is real money, so we default to one run and flag its uncertainty instead.
On the Chinese platforms: completely fair. 36Kr isn't an A/B, it's PR; Juejin and Zhihu are the actual self-publish A/B we're missing. And since both wrap or nofollow outbound links, they're recall plays, not backlink plays — worth stating as such so nobody expects the same post to do SEO work at the same time.
All three land. Editing the recommendation layer to be per-model now.
On cost: you may not need N runs weekly, only a noise floor. The six runs from Aug 19–24 already in your logs are a free calibration set — measure per-model variance once, store it, then keep the weekly cadence at one run and read it against that floor. N becomes an occasional recalibration cost instead of a recurring multiplier, and the card can say "blank, but this model blanks in 2 of 6 runs" instead of a uniform ±10%.
The variance won't be uniform either — Perplexity at 5/5 and Qwen flipping inside a day are different regimes — so you only need real N on the unstable cells. And worth logging "returned nothing at all" separately from "answered but didn't mention the brand": the first is usually infra, and marking those invalid saves re-runs you'd otherwise pay for.
The "your own homepage is the weakest citation you can have" point lines up with something I ran into from the pre-launch side — no backlinks yet, so anything an AI cites about my site right now can only be self-sourced by definition, which is apparently the least trusted category. Makes the "give it a second independent source" advice more urgent than I'd been treating it, since I don't have the luxury of "good enough for now."
The Qwen zero-mentions finding is the one I'd want to dig into more — is that purely a training-data gap (nothing in Chinese to learn from), or does it also reflect a difference in what counts as a citable source for that ecosystem specifically? If it's the latter, "post once in Chinese" might not close the gap the way "post once on Dev.to" does for the Western models, since the trust bar might be structurally different, not just a language one.
Great questions — both point at the same limit: we can see whether an AI cites you, but not yet why it doesn't.
On self-sourcing: agreed, and it's exactly why "second source" was the highest-leverage fix in the report. Pre-launch you're not helpless, though — the cheapest citations that aren't your own domain: a GitHub repo README, one Dev.to writeup, a Product Hunt upcoming page, a directory entry (Toolify, TAAFT, Futurepedia — or Chinese ones like ai-bot.cn). Each is an afternoon of work.
On Qwen: honestly, we don't know yet whether it's a training-data gap or a different bar for what counts as citable. Qwen is trained Chinese-first, so the language gap is nearly certain. But which Chinese sources its search layer trusts (Baidu Baike, Zhihu, CSDN, WeChat articles) is a different system than the English web — likely structural, not just linguistic. The test is cheap: publish one Chinese post, re-run the same questions against Qwen, see if a citation appears. That kind of before/after is exactly what we built. If you run it, I'd genuinely want to see the result.
That list is exactly the kind of concrete, afternoon-sized action I needed instead of another abstract "build authority" suggestion. Dev.to writeup and a Product Hunt upcoming page are both things I can just go do this week — no excuse to keep treating this as theoretical when the fix is that cheap.
The Qwen test is genuinely tempting to actually run, mostly because "we don't know yet" from someone who built a tool specifically to measure this is a more honest answer than most people would give, and it means there's real, unclaimed territory here rather than a solved problem I'd just be repeating. Don't have a Chinese-language post ready, but if I put one together I'll run the before/after and report back what actually happens to the Qwen citations — genuinely curious myself now, not just being polite about it.
This experiment highlights both the immense promise and the inherent instability of AI-driven discovery, showing that while major LLMs can serve as high-converting recommendation engines, hallucination risks remain a constant wildcard. As AI Engine Optimization (AEO) becomes crucial for modern products, controlling the narrative across training datasets will be just as vital as securing high search rankings to ensure brand accuracy.
Exactly — and the report has a concrete example of the "train the dataset narrative" problem: when Perplexity did cite the product, the only citation was the product's own site — i.e. the least-independent source. So it's not just that the narrative controls rank; the sources that anchor that narrative are themselves a ranking input. Controlling both is the real AEO problem.
Are you working on AI-related discovery too, or more on the search/ops side?
Trustworthiness is a measurement system designed into the model. Perplexity's rule - 'independent verification = trustworthy' - creates a discovery gap that punishes solo founders. Your own homepage is zero-trust because it has obvious incentive alignment. Dev.to/PH/directories are higher-trust because they imply editorial gate-keeping. The gap between 'mentioned in training data' and 'recommended in output' reveals what the model actually considers verified. Qwen's zero knowledge of English products isn't a Qwen problem - it's a selection bias in your sources. Chinese AIs trained on Chinese sources measure prestige through Chinese citations. That's not a bug, it's their measurement system working correctly for their audience. The leverage move isn't 'get mentions' - it's 'understand which measurement system controls visibility in each AI ecosystem.'
This is the sharpest framing of the report yet — "trustworthiness as a measurement system designed into the model" is exactly what we keep running into, and you've articulated the implication we hadn't: the source-selection bias isn't a Qwen problem, it's Qwen's measurement system working correctly for its audience.
That reframe has a practical consequence for what we're building: optimizing for "mentions" treats the symptom. The real variable is which measurement system controls visibility in each ecosystem — and they don't converge. An English-first brand's trust graph looks nothing like a Chinese-first brand's. So the product question stops being "how do I get cited" and becomes "which ecosystem's measurement system is my bottleneck, and what are its specific trust signals?"
One thing I'd genuinely like your take on: do you think these measurement systems will converge over time — the way PageRank became the de-facto standard for "the web" — or stay fragmented per ecosystem? We're betting on fragmented; it's why we monitor 6 engines as separate channels rather than one blended score.
(As a side note, the report's Perplexity example is a clean illustration of your point: it did cite the product — and the only citation category it trusted was the product's own site, the least editorially-gated source. The model's "verified" bar and ours don't always align.)
Interesting that independent verification mattered more than the product’s own site.
Did adding one strong third-party source materially change the AI recommendations afterward?
Honest answer: we haven't systematically run that experiment yet. The report was a single snapshot — we measured what each engine cited at one point in time, but we didn't add a third-party source mid-study and re-run the same questions.
That before/after is exactly the test we want to run — and Beryxa is the perfect candidate since we already have your baseline report. If you publish one independent writeup about Beryxa (a Dev.to post, a directory listing, a launch story anywhere off your own domain), we re-run the same 6 questions and measure whether the citations — and the recommendation — actually move.
That's the experiment that tells us whether "add a second source" is a real lever or just a theory.