hreflang said English, canonical said Chinese. Google believed the canonical.
I'm Cody Wang. I build and optimise Prismic, SvelteKit and Next.js sites, mostly for clients in New Zealand and for Chinese brands going international.
Last week I put a small SEO audit tool online. Most of the week went into the unglamorous half of launching: making sure search engines could actually read what I'd built. I found one bug that had quietly removed half the site from the index, and it took me three days to spot because every single page looked fine on its own.
#What the site is
It's a detector. You paste a URL, it tells you what the site is built with and lists the fixable problems.
Two languages, one URL, language in a query parameter:
/tools/cloudflare Chinese (default)
/tools/cloudflare?lang=en English
66 URLs, all server rendered. The content is in the HTML, no JavaScript required.
#What I was telling Google
Here's the head of the English page for the Cloudflare detector:
<html lang="en">
<title>Cloudflare Detector — check if a website uses Cloudflare</title>
<link rel="canonical" href="https://prismicaudit.com/tools/cloudflare">
<link rel="alternate" hreflang="zh-Hans" href="https://prismicaudit.com/tools/cloudflare?lang=zh">
<link rel="alternate" hreflang="en" href="https://prismicaudit.com/tools/cloudflare?lang=en">
<link rel="alternate" hreflang="x-default" href="https://prismicaudit.com/tools/cloudflare">
Look at the canonical. The English page is declaring that its real URL is the Chinese one.
So I was saying two contradictory things on the same page. hreflang says the English version lives at ?lang=en. The canonical says the canonical version lives at the bare URL, which serves Chinese. When those two disagree, Google follows the canonical. The English page becomes a duplicate of the Chinese page, and 1,600 words of English copy I wrote for English searchers gets folded into a page in a language those searchers can't read.
The cause was dull. My head builder took lang as its first argument and used it for the <html> attribute, the title, the description, the structured data, the font stack. Then it built the canonical out of the path alone, because that line was written before the language switching existed and nobody updated it when the language switching arrived. One argument that got passed in and never used, on every page of the site.
#How I noticed
Not from a dashboard alert. There isn't one for this.
Search Console showed 199 impressions and 8 clicks across four weeks. Four of the five pages that earned a click were ?lang=en URLs. The Chinese pages, which were the canonical targets, were picking up scraps.
That's the tell. Almost everything getting impressions was something I'd told Google to ignore.
You can check this on your own site faster with curl than with any tool:
curl -s "https://prismicaudit.com/tools/cloudflare?lang=en" \
| grep -o '<link rel="canonical"[^>]*>'
If the canonical that comes back has no lang in it while the page itself is in English, you have the same problem.
#The fix
Each language version gets a canonical pointing at itself:
<link rel="canonical" href="https://prismicaudit.com/tools/cloudflare"> <!-- zh -->
<link rel="canonical" href="https://prismicaudit.com/tools/cloudflare?lang=en"> <!-- en -->
The hreflang block stays, except x-default now points at the English version. That one is a judgement call and not a rule. x-default is the page you show when no language matches, so it should be the version your audience actually searches in. Mine is English, so that's where it points. On a site aimed at the Chinese market I'd point it the other way.
The sitemap needed the same treatment. It was emitting one <loc> per path with three alternates, which never put the English URL in the sitemap at all. Now each language gets its own entry, so 66 URLs became 132.
#Verifying it
I don't trust "looks right." I wrote a loop that walks every URL in the sitemap and checks the canonical on the English version:
U="https://prismicaudit.com"
curl -s "$U/sitemap.xml" | grep -o '<loc>[^<]*</loc>' \
| sed -e 's/<loc>//' -e 's/<\/loc>//' | grep -v 'lang=en' | sed "s#^$U##" > /tmp/paths.txt
while read -r p; do
got=$(curl -s "$U$p?lang=en" | grep -o 'rel="canonical" href="[^"]*"' | head -1)
[ "$got" = "rel=\"canonical\" href=\"$U$p?lang=en\"" ] || echo "FAIL $p -> $got"
done < /tmp/paths.txt
66 paths, 66 passes, no output. Mobile Lighthouse on a detector page still reads 100 for performance, accessibility, best practices and SEO. The fix didn't cost anything measurable.
#What I got wrong while fixing it
The homepage had its own head builder, separate from the page template, so the same bug needed fixing twice. I patched the homepage by running a regex over the already-rendered HTML instead of fixing the template that produced it.
It passes every check I have. It's also fragile in a way I'll pay for later: if the order of those three link tags ever changes, the regex matches nothing, silently, and the bug comes back with no error anywhere. I left a note in the file. Next time I touch that template it goes in the template.
Two head builders and one regex. The actual lesson wasn't about hreflang.
#What I still don't know
Whether it worked.
Canonical and hreflang are hints for crawlers. I've changed the hints. Now Google has to come back, re-crawl 132 URLs, and re-evaluate which page is the real one. That takes days to weeks, and until it does, both versions are in a weird state where the signals contradict what Google last saw.
So I'm in the waiting part. Nothing to design, nothing to ship, just checking Search Console every few days to see whether the English URLs start collecting impressions the way they should have been all along.
If you run a multilingual site, the curl loop takes ten minutes. Nothing in a Lighthouse report will ever show you this one.