My href regex found 0 links on a page with 101 of them

5 min read Web ScrapingPythonRegexHTMLDebugging

A link-extraction regex returned nothing on a page full of links. The page was fine. It ships minified HTML, where quotes around an attribute value are optional, so href=/media/x.pdf matches nothing that expects href="...".

TL;DR · THE FIX

Attribute quotes are optional in HTML, and minifiers drop them. Every href="([^"]+)" pattern you have ever copied silently returns zero on a minified page. Match all three forms, href=(?:"([^"]+)"|'([^']+)'|([^\s>]+)), resolve with urljoin, and drop any result whose hostname is localhost. When your own tool returns zero, check the tool before you draw conclusions about the page.

The symptom

I run a small watcher that reads a school’s parent-letters page and tells me when a new letter appears. It kept reporting that there was nothing readable: the dates it wants live inside linked PDFs, and it was finding no PDFs to open.

So I opened the page myself. It is a wall of letters going back years, every one of them a link.

Checking by hand, with the one-liner everybody reaches for:

curl -s https://example.school/eltern/elternbriefe | grep -oiE 'href="[^"]*\.pdf"'

Nothing. No output at all.

Zero is a convincing answer. It does not look like a broken tool, it looks like a fact about the page: this page links no PDFs. So I went off looking for reasons the page might be hiding them from me. A login wall. Markup drawn by JavaScript after load. A bot check serving me something different from what the browser gets.

None of that was true. I spent the time debugging a page that was never broken.

The cause

The page ships minified HTML, and a minifier drops everything it is allowed to drop. Quotes around an attribute value are optional in HTML when the value contains no spaces, quotes, or a handful of other characters. So the markup that actually arrives looks like this:

<li class=item><a href=/media/2026/09/letter-14.pdf class=dl>Elternbrief</a></li>

Not one attribute on the page is quoted. Every regex, tutorial and half-remembered one-liner for pulling links out of HTML assumes href="...", because that is how humans write HTML and therefore how every example is written.

Measured on the live page, both patterns over the same bytes:

result
page size34,913 bytes, 50 lines
href attributes present128
…of those, quoted0
href="([^"]*\.pdf)"0 matches
three-branch pattern (below)101 PDF links

The pattern was measuring my assumption about how pages are written, and the assumption was wrong.

The fix

The repair is in the pattern, since I cannot change the site. Accept all three forms an attribute value can take: double quoted, single quoted, and bare.

import re
import urllib.parse

PDF_LINK_RE = re.compile(r"""href=(?:"([^"]+)"|'([^']+)'|([^\s>]+))""", re.I)

def pdf_links_in(raw, page_url):
    out = []
    for m in PDF_LINK_RE.finditer(raw.decode("utf-8", "replace")):
        href = (m.group(1) or m.group(2) or m.group(3) or "").strip()
        if not href:
            continue
        href = urllib.parse.urljoin(page_url, href)          # not string concat
        if not href.lower().split("?")[0].endswith(".pdf"):
            continue
        if not href.startswith(("http://", "https://")):
            continue
        host = urllib.parse.urlsplit(href).hostname or ""
        if host in ("localhost", "127.0.0.1"):               # see below
            continue
        out.append(href)
    return list(dict.fromkeys(out))                          # dedupe, keep order

Three details in there each cost me something.

Keep all three branches. Do not “simplify” this back to the quoted form once it works. The bare-value branch is the reason it works here, and it is also the branch that looks redundant to anyone reading it later.

Use urljoin rather than string concatenation. The hrefs are a mix of absolute URLs on another host and root-relative paths. Gluing a base onto a path that is already absolute produces nonsense that still looks like a URL.

Drop localhost. One of those 101 links points at http://localhost:8080/...: the CMS leaked its own preview host into published markup. Fetch that and your scraper hangs waiting for a machine that is not on the internet, which then reads as “the site is slow” rather than “this link is unfetchable by construction”.

The better fix, if you can take it

Use a real parser. selectolax, lxml.html, or BeautifulSoup all handle unquoted attributes correctly, because they implement the HTML spec instead of the subset of HTML that appears in examples:

from selectolax.parser import HTMLParser
hrefs = [n.attributes.get("href") for n in HTMLParser(html).css("a[href]")]

The regex above exists in my watcher because that script has no dependencies by design and I wanted to keep it that way. That is a legitimate trade, and the cost is that you own every corner of the HTML grammar you did not think about. Attribute quoting is one of the cheapest corners to get wrong.

One more trap in the same job

Once the links were flowing, one letter in twelve extracted zero characters of text. Those are scans: an image of a page inside a PDF wrapper, where pypdf correctly reports that there is no text, because there is none.

“I read it and found nothing” looks identical to “there was nothing to find” unless you treat it as a separate case:

text = "\n".join((p.extract_text() or "") for p in PdfReader(tmp).pages)
if len(text.strip()) < 40:
    status = "unreadable (probably a scan)"   # not "no dates found"

Same bug as the main one: an empty result read as information about the world when it was information about the instrument.

The part worth keeping

The measurement that mattered here took ten seconds and I did it last instead of first: look at the bytes the server actually sent. If the whole document is one dense line, you already know how much of it will match a pattern written against tidy example HTML.

When a scrape comes back empty, suspect your pattern before you suspect the site. A login wall, JavaScript rendering and bot detection are all real, and they are all more interesting than the boring answer, which is why they are the theories you reach for first. A zero from your own instrument feels like data, and there is nothing on screen to disagree with, so it is the easiest thing to be confidently wrong about.

Related fixes

Discussion

Powered by GitHub. Sign in to leave a comment.