Published on

Scraping: fail loudly when the page layout changes

Authors
On this page

The worst scraper bug is the quiet one. A site changes its layout, your selector matches nothing, the loop runs over zero items, and the job exits 0. Nobody notices until someone asks why a month of data is empty.

Two small habits catch this.

1. Compare against a number the page gives you

Pages often say how many items they have: a result count, numberOfItems in structured data, a "showing 1-20 of 340" label. Parse that too, and compare:

class Drift(Exception):
    pass

def parse(h):
    items = [{"name": n, "price": float(p)} for n, p in re.findall(
        r'<span class="name">(.*?)</span><span class="price-tag">\$([\d.]+)</span>', h)]
    if not items:
        raise Drift("0 items parsed")
    expected = int(re.search(r'"numberOfItems": (\d+)', h).group(1))
    if len(items) != expected:
        raise Drift(f"parsed {len(items)} items, page says {expected}")
    return items

Zero results should fail the job. So should a mismatch with the page's own count.

2. Keep a saved copy of a good page

Save one real page as a fixture, and run the parser on it in your tests. That tells you the parser works. It cannot tell you the live site is unchanged, so also run the same check on every live fetch.

What it printed

I saved a good page, then changed the live layout the way a redesign would:

fixture parses: [{'name': 'Widget 1', 'price': 10.5}, {'name': 'Widget 2', 'price': 11.5}, {'name': 'Widget 3', 'price': 12.5}]
live page after redesign -> Drift: 0 items parsed

The fixture still parsed. The live page raised Drift at once, with a message that names the problem.

What to do with the error

  • Let it fail the job, so your scheduler or monitor sees a non-zero exit.
  • Do not write partial results over good ones.
  • Log the URL and the count, and keep the bad HTML for debugging.

A scraper that stops is annoying. A scraper that is wrong and silent costs you the data you thought you had.