Search Console: "Page with redirect" on 106 pages that all return 200
An index-failure email listed 102 URLs as Page with redirect, and every one of them loaded fine in a browser. Astro's directory build format emits /about/ in the sitemap and rel=canonical, every internal href in the source said /about, and the host 308s between them. Googlebot only ever reached the site through a redirect.
TL;DR · THE FIX
If your sitemap says /about/ and your links say /about, the host's 308 hides the mismatch from every human and from nobody at Google. Fix every internal link to the canonical form, including the ones that are not in a template: markdown body links, href helpers, hardcoded hub lists, server-rendered demos, and any build-time JSON search index. Measured on dist/: no-slash internal links 12 to 5 to 0, and 136/136 sitemap entries matching.
The symptom
An email from Search Console said indexing validation had failed. The report listed 106 pages under “Page with redirect” and enumerated 102 of the URLs.
Every one of them returned 200 in a browser. There was no redirect chain, no stale URL from an old structure, no www versus apex problem. These were live, linked, correctly rendering pages, and Google was calling them redirects.
I pasted one into the URL Inspection tool first. It confirmed the verdict and said almost nothing about the cause, which makes sense from Google’s side: all it knows is that the URL it crawled was not the URL it ended up at.
What made this hard to see
The site is Astro on Cloudflare Pages with the default build.format: 'directory', which writes about/index.html rather than about.html. Two things follow from that, and both are correct behaviour:
- The sitemap integration and
rel=canonicalboth emit the directory form,https://example.dev/about/. - Cloudflare Pages serves
about/index.htmlat/about/and 308s/aboutto/about/, so the no-slash form still works.
Meanwhile every internal link in the source had been written the short way:
<a href="/about">About</a>
<a href="/fixes">All fixes</a>
So the site worked. A human clicks /about, gets a 308, lands on /about/, and never notices. No broken link, no 404, no console error, nothing in a Lighthouse run.
Googlebot’s experience differs in one way. It discovers URLs by following links, so almost every URL it discovered was the no-slash form. It followed the redirect, arrived at the slash form, and read a rel=canonical pointing at the slash form. From its point of view the discovered URL is a redirect to the canonical, and that is precisely what “Page with redirect” means. Search Console was describing, accurately, a site whose links and canonical form disagreed.
The count fits that reading too. The pages Google had indexed cleanly were the ones it reached from the sitemap rather than from a link.
What I checked before believing it
Trailing-slash theories are cheap and sound plausible, so I wanted the site to prove it. The place to check is the build output, since the source can contain a link the build rewrites and a template can contain a link no page ever renders. I wrote a script over dist/ that counts internal href values that neither end in / nor point at a file or anchor.
It found 12. Fixing the templates brought it to 5, and those five were in places where nobody looks for a link:
- Markdown body links. Prose in content files, written by hand months apart, all pointing at
/fixes/some-slugwith no trailing slash. - A
tagHref()helper. One function, one missing character, every tag page on the site. - A hardcoded hub list. An array of
{ title, href }objects for a section index, sitting in a component rather than in content. - A server-rendered demo. A terminal-style hero that prints paths as output. They look like sample text and they are real anchors.
- A build-time JSON search index. The client-side search reads a generated JSON file of slugs and turns them into links at runtime, so these bad URLs lived in a build artifact that no HTML scan touches. This was the nastiest of the five.
The search index is the general shape of the trap. Anywhere a path is stored as data instead of written as markup is invisible to a scan of the rendered pages: a JSON-LD block, an OG URL, a redirect map, an RSS feed, an llms.txt.
The fix
Make the internal links match the canonical form your build format produces. With Astro’s directory format that means trailing slashes everywhere, data files included.
- <a href="/about">About</a>
+ <a href="/about/">About</a>
- const tagHref = (tag) => `/tags/${slugify(tag)}`;
+ const tagHref = (tag) => `/tags/${slugify(tag)}/`;
Going the other way, build.format: 'file' with no slashes anywhere, is equally fine. What does not work is one convention in the sitemap and another in the links, because the host’s redirect makes that combination look healthy.
The measurement
Three numbers, all taken from dist/ after a real build:
- Internal no-slash
hrefcount: 12 to 5 to 0. - Search-index slugs in no-slash form: 0.
- Sitemap entries in canonical form: 136 of 136.
Then the live check, since a green build says nothing about the deployed site: production returned 200 on /, /fixes/, /tags/csv/, /about/ and /sitemap-0.xml, and the rendered HTML carried slashed links.
Keep it fixed
A link like this comes back the first time somebody types a href from memory, and it will look fine when they do, because the redirect still works. So the fix is only finished once something fails the build over it:
// scripts/check-links.mjs, run after astro build
import { readdirSync, readFileSync, statSync } from "node:fs";
import { join } from "node:path";
const walk = (dir) =>
readdirSync(dir).flatMap((f) => {
const p = join(dir, f);
return statSync(p).isDirectory() ? walk(p) : [p];
});
// Code is prose about links, not links: <pre> for fenced blocks,
// inline <code> for a path mentioned mid-sentence.
const stripCode = (html) =>
html.replace(/<pre[\s\S]*?<\/pre>/g, "").replace(/<code[\s\S]*?<\/code>/g, "");
const bad = [];
for (const file of walk("dist").filter((f) => /\.(html|json)$/.test(f))) {
const raw = readFileSync(file, "utf8");
const text = file.endsWith(".html") ? stripCode(raw) : raw;
for (const [, url] of text.matchAll(/["'](\/[^"'#?\s]*)["']/g)) {
if (url === "/" || url.endsWith("/")) continue;
if (/\.[a-z0-9]{2,5}$/i.test(url)) continue; // real files
bad.push(file + ": " + url);
}
}
if (bad.length) {
console.error(bad.length + " internal links missing a trailing slash:");
for (const b of bad) console.error(" " + b);
process.exit(1);
}
console.log("All internal links use the canonical trailing-slash form.");
Two details are worth copying. Scanning .json as well as .html is what catches the search index, which had survived two rounds of fixing. Stripping code out first, fenced blocks and inline spans both, stops the check failing on its own documentation: the first run flagged exactly one link, the href="/about" in the code sample above, on the page you are reading. A check that fires on prose about the bug gets switched off within a week.
Wire it into the build so it is not a thing to remember:
"build": "astro build && node scripts/check-links.mjs"
Then break something on purpose. Inject a no-slash link into a built page, confirm the script exits 1 and names the file, remove it, confirm the count returns to 0. Until you have watched a guard fire, you do not know that it can.
The lesson
Every no-slash link on that site worked for a year, so nothing in the ordinary development loop could surface the problem: clicking around, running a link checker, reading a Lighthouse report all passed. The redirect that kept everything working was also what hid the mismatch. The only consumer that cared was the one that treats “the URL I found” and “the URL I landed on” as two separate facts.
If you are on Astro and Pages, your build format decides your canonical URL shape and your links have to agree with it. Check the built output rather than the source, and check the JSON in the built output too.
Discussion
Powered by GitHub. Sign in to leave a comment.