Link preview blank? Open Graph tags set by JavaScript never reach the crawler
Social crawlers read the HTML bytes and do not run your JavaScript, so og:title and og:image written client-side never reach them, even while DevTools shows the right tags. Check with curl and a crawler user-agent, then inject the tags server-side, here with a Cloudflare Pages Function.
TL;DR · THE FIX
If your meta tags are written by client-side JavaScript, no social crawler will ever see them, because no social crawler runs your JavaScript. DevTools hides this from you: the Elements panel shows the DOM after scripts have run, while the crawler reads the response bytes before they have. Check it with curl and a crawler user-agent, and inject the tags server-side.
The symptom
I pasted a share link into a chat and got back a grey box: no image, a generic title, and a one-line description that was true of every share link on the site rather than of this trip. The page itself was fine. It opened, fetched the trip, rendered the collage, and the browser tab said the right thing. The page worked and the preview of the page did not.
What I tried first
My first guess was that the image was private. The collage is served from storage, and an expired signed URL or a bucket that is not public would give a crawler nothing to draw. It opened in a logged-out browser with no token and no session, so that was ruled out.
My second guess was a cached preview. Every platform caches scrapes aggressively, and the first time that URL was ever shared the page had no content on it, so a stale scrape from that moment would look exactly like this. A fresh scrape came back with the identical empty card.
Both guesses were theories about the crawler, and I had not yet looked at what the crawler was handed.
What was happening
So I stopped opening the page in my browser and asked for it the way a crawler does: one request, same URL, the crawler’s user-agent, then read the bytes.
curl -s -A "facebookexternalhit/1.1" \
"https://example.pages.dev/share/?token=..." \
| grep -E '<meta (property="og:|name="twitter:)'
The preview title was still the placeholder, and there was no og:image in the response at all.
The page is a static file. Its <head> ships with a generic title, and after the fetch resolves this runs:
// share/index.html
document.getElementById("og-title").setAttribute("content", `${data.group_name} - Example`);
document.title = `${data.group_name} - Example`;
That is the only place the real title is ever set, and it runs in a browser. A crawler reads the response and leaves without running your JavaScript.
Two things were wrong, and they are the same mistake. The one preview tag anybody had tried to maintain lived in the exact place a crawler cannot see, and og:image, the tag that decides whether a card looks like anything at all, was never written anywhere.
DevTools hid it. The Elements panel shows the page after JavaScript has run, so the og:title there read correctly, with the real name in it, every time I looked. view-source shows the bytes. Two panels in the same window disagreed, and only one of them was what got sent. The document.title line next to it made this worse, because that half works: the browser tab really does update, so nobody suspected the line above it.
The fix
Put the tags in the response instead of in the browser. On Cloudflare Pages that is a function on the same route, in front of the static file:
// functions/share/index.js
const CRAWLERS = [
"facebookexternalhit", "twitterbot", "linkedinbot", "whatsapp",
"slackbot", "telegrambot", "discordbot", "googlebot", "bingbot",
"applebot", "imessage", "icloud",
];
const isCrawler = (ua) =>
!!ua && CRAWLERS.some((bot) => ua.toLowerCase().includes(bot));
const esc = (s) => String(s)
.replace(/&/g, "&").replace(/"/g, """)
.replace(/</g, "<").replace(/>/g, ">");
export async function onRequest(context) {
const url = new URL(context.request.url);
const token = url.searchParams.get("token");
const ua = context.request.headers.get("user-agent") ?? "";
// Not a crawler: hand back the static file, untouched.
if (!isCrawler(ua) || !token) return context.next();
let title = "Trip Story - Example";
let description = "See the shared memories from this trip.";
let image = "";
try {
const res = await fetch(`${API}/share`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ share_token: token }),
});
if (res.ok) {
const data = await res.json();
if (data.group_name) {
title = `${data.group_name} - Example`;
description = `${data.photo_count} photos · ${data.member_count} members`;
image = data.image_url ?? "";
}
}
} catch { /* fall through to the generic meta */ }
const staticRes = await context.next();
let html = await staticRes.text();
const tags = [
`<meta property="og:title" content="${esc(title)}" />`,
`<meta property="og:description" content="${esc(description)}" />`,
`<meta property="og:image" content="${esc(image)}" />`,
`<meta property="og:image:width" content="1204" />`,
`<meta property="og:image:height" content="904" />`,
`<meta name="twitter:card" content="summary_large_image" />`,
`<meta name="twitter:title" content="${esc(title)}" />`,
`<meta name="twitter:description" content="${esc(description)}" />`,
`<meta name="twitter:image" content="${esc(image)}" />`,
].join("\n ");
html = html.replace("</head>", ` ${tags}\n</head>`);
return new Response(html, {
status: staticRes.status,
headers: { "Content-Type": "text/html;charset=UTF-8" },
});
}
Four things in there are deliberate. context.next() on the non-crawler path returns before anything else happens, so real users get the static asset off the edge with no upstream fetch and no HTML rewrite in the way; the function is invisible to everybody it is not for. The upstream fetch is wrapped in try with the fallback values already assigned, so if the API is down the crawler gets a generic card rather than a 500, and a platform that scrapes during an outage does not cache an error page against your URL. Every injected value goes through esc, because these strings come from user-supplied content and land inside a double-quoted HTML attribute, where one unescaped " ends the attribute early and the rest of the tag becomes markup. And og:image:width and og:image:height are set, because several platforms will not render the large-card layout until they have fetched and measured the image, and some give up first; declaring the dimensions gets you the big card on the first scrape instead of the third.
Verify it the same way you found it
The check is the same two requests, differing only in the header. Measured against the live route:
U="https://example.pages.dev/share/?token=..."
curl -s -A "facebookexternalhit/1.1" "$U" \
| grep -cE '<meta (property="og:|name="twitter:)'
# 13
curl -s -A "Mozilla/5.0 ... Chrome/127" "$U" \
| grep -cE '<meta (property="og:|name="twitter:)'
# 4
Thirteen tags including og:image for the crawler, four and no image for the browser, from one URL. It is also the only form of proof that would have caught the bug in the first place.
The platform debuggers are worth a pass too, since they force a re-scrape and the old empty card is cached: Facebook’s Sharing Debugger, LinkedIn’s Post Inspector, and X’s card validator.
The lesson
When something is broken for a client you do not control, your own browser is a different client running different code, and here it was the one client that could not see the problem. Make the request the way that client makes it and read what it got back. One curl with the right User-Agent would have ended this in a minute, at any point in the two days I spent inspecting a DOM that was never sent to anyone.
Discussion
Powered by GitHub. Sign in to leave a comment.