My Cloudflare Pages deploy returned success and served 500 on every route

4 min read CloudflareCloudflare PagesDeployment

A direct upload to Cloudflare Pages answered success: true, handed back a deployment URL, and every single path on it returned HTTP 500 - including the live domain, because a deploy with no branch named goes to production. The manifest is not the files. Here is the three-step upload that actually works.

TL;DR · THE FIX

A Pages deployment manifest is a map of {path: content-hash}. The content behind those hashes has to be registered first, in a separate call to /pages/assets/upload, using wrangler's own key: blake3(content + extension) truncated to 32 hex chars. Post the manifest on its own and the API happily creates a deployment whose every path points at content it has never seen - a valid deployment of nothing. It reports success because your request really was accepted. Deploy to a throwaway preview branch and fetch a route off it before you go near production.

The symptom

I deployed my site with a direct upload to Cloudflare Pages. The API answered:

{ "success": true, "result": { "url": "https://a3f0e1c8.unstuck.pages.dev", "id": "..." } }

Then I opened it. Every route returned HTTP 500: the homepage, the index pages, the individual posts, all of them.

And it was the live domain. A Pages direct upload with no branch specified goes to the project’s production branch, so an upload shape I had never run before went straight onto the apex on its first attempt.

What I tried first

My first guess was a broken build. The same dist/ folder served correctly from a local static server, and every file in it was present and well-formed.

My second guess was that Cloudflare was having a bad afternoon, and ruling that out is what pointed at the upload. I rolled back to the previous deployment from the Pages dashboard and every route came back within a minute or two. The platform, the project and the domain were all fine. The only broken thing was the deployment I had just made.

Then the measurement that settled it. I ran the same upload again, naming a throwaway preview branch:

python deploy_cf.py probe-deploy
preview     https://probe-deploy.unstuck.pages.dev/  ->  500
production  https://unstucked.dev/                    ->  200

Same folder, same code, same command, two targets, at the same moment. The preview served 500s while production, untouched, served 200s, so the upload was the problem.

What was happening

My uploader posted one request: a deployment carrying a manifest plus the files as multipart parts. I had copied that shape from another project of mine that deploys to a different kind of target, it looked reasonable, and the API accepted it.

A Pages direct upload is three calls. GET /accounts/<acct>/pages/projects/<project>/upload-token returns a short-lived JWT. POST https://api.cloudflare.com/client/v4/pages/assets/upload, with that JWT, registers the file contents in batches, each entry keyed by a content hash. Only then does POST /accounts/<acct>/pages/projects/<project>/deployments post the manifest, and a manifest is a map:

{
  "/index.html": "a3f0e1c8...",
  "/fixes/index.html": "9c21b7d4..."
}

Paths on the left, content hashes on the right. The deployment call says “serve these paths, and here is which blob each one is”. The asset upload is what puts the blobs where Pages can find them.

I had done the deployment call and nothing else. Every hash in my manifest was correctly computed and pointed at content the platform had never received. The result is a structurally valid deployment of nothing, which from the outside looks exactly like a 500 on every path.

The success flag was not lying

The API was asked to accept a deployment request. The request was well-formed, so it accepted it and created a deployment, and success: true reports that accurately. Whether the resulting site serves anything was never the question that call answered. It could not be: at the moment it returns, nothing has fetched a page.

The fix

Do the first two calls, and use wrangler’s own key so the hashes you claim are the hashes Pages computes:

from blake3 import blake3   # pip install blake3

def asset_key(data: bytes, path: str) -> str:
    """wrangler's asset hash: blake3 over the content PLUS the extension, 32 hex."""
    ext = os.path.splitext(path)[1].lstrip(".")
    return blake3(data + ext.encode()).hexdigest()[:32]

Register the content under those keys:

r = requests.get(f"{API}/accounts/{ACCOUNT}/pages/projects/{PROJECT}/upload-token",
                 headers={"Authorization": f"Bearer {TOKEN}"})
jwt = r.json()["result"]["jwt"]

payloads = [{
    "key": key,
    "value": base64.b64encode(data).decode(),
    "metadata": {"contentType": mimetypes.guess_type(full)[0] or "application/octet-stream"},
    "base64": True,
} for rel, (key, data, full) in files.items()]

# batch them: keep each request comfortably under ~15 MB of base64
resp = requests.post(f"{API}/pages/assets/upload",
                     headers={"Authorization": f"Bearer {jwt}",
                              "Content-Type": "application/json"},
                     data=json.dumps(batch))

The asset upload goes to the account-less /pages/assets/upload and authenticates with the JWT from the first call rather than with your API token. Only then post the manifest:

manifest = {rel: key for rel, (key, _, _) in files.items()}
requests.post(f"{API}/accounts/{ACCOUNT}/pages/projects/{PROJECT}/deployments",
              headers={"Authorization": f"Bearer {TOKEN}"},
              files=[("manifest", (None, json.dumps(manifest), "application/json"))],
              data={"branch": branch})

Deduplicate by key before uploading. A static build has plenty of byte-identical files, and there is no reason to send the same blob twice.

The part worth keeping

The bug was the missing call. The outage came from something else: I never chose production. I left the branch argument off, the default chose production for me, and the first run of a shape I had never tested landed on the live domain. Every step of that is unremarkable on its own, which is why it works.

Two habits come out of it. Anything that ships has to be aimed. If an operation has a target and you did not name one, you have delegated the decision to a default chosen by someone optimising for a different situation than yours, so name the branch every time, including when it is the one you wanted anyway. And a success response describes your request and says nothing about the artifact, so finish the job by asking the artifact:

curl -si https://probe-deploy.unstuck.pages.dev/ | head -1
# HTTP/2 200

Fetch a real route off the preview URL, read the status line, and only then deploy to production. It costs one command, and it is the only step in the sequence that checks the thing you care about.

Related fixes

Discussion

Powered by GitHub. Sign in to leave a comment.