laya-scraper GitHub

Open source, MIT

An AI writes the scraper once. Plain code runs it after that.

Give it a website and one line of plain English: "grants", "patents", "datasets", "policy papers". It finds the listing, pages through it, designs a schema and extracts every item into a table. The AI is used only to write a small recipe for that site. Every run after is plain code: no AI, no cost, the same answers each time.

harvest.py
$ python harvest.py https://patents.justia.com/ "patents"
plan:   fields: title, patent_number, inventors,
        assignee, filing_date ...
locate: LLM picked #3 /company/linkedin
        (lists actual patents, paginated)
paging: next-link, 100 items
  round 0: title ok, patent_number FAIL,
           inventors ok, assignee FAIL ...
  round 1: title ok, patent_number ok,
           inventors ok, assignee ok ...
    patent_number validated (4/4 correct)
    assignee      validated (4/4 correct)
filled: title 100/100, patent_number 100/100,
        inventors 100/100, assignee 100/100 ...
  • 25LLM calls to build a scraper for a site, once
  • $0.069what that build cost on ukri.org, measured
  • $0every run after, because no AI is involved
  • 78sto pull 100 items on a re-run

Measured on ukri.org with the prompt "research funding calls", priced with MiniMax-M3 at $0.30 in and $1.20 out per million tokens. Messier sites take more calls.

Watch one real run, start to finish

Six narrated minutes: the prompt becomes a plan, the crawler finds the listing, selectors are graded and repaired, and the same recipe runs again with no AI at all.

720p, 8 MB, 6 min 24 s. There is also a 1080p version at 42 MB.

The AI writes the scraper. Plain code runs it.

The LLM never scrapes anything. On the first run it writes a recipe: a small JSON file holding the listing URL, how to page through it, and one proven selector per field. Every later run follows that recipe.

Diagram: the AI steps, plan, locate and map, happen once and produce a recipe; the plain code steps, page, fetch and apply selectors, run every time.

An LLM makes a handful of decisions

It turns your prompt into a plan and a field schema, picks the real listing among the crawler's candidates, maps each field to a page element, and grades its own choices. MiniMax by default, or any OpenAI or Anthropic compatible endpoint, including a local model.

laya does the per-page judging

A small non-generative decision model on the GPU: is this page about the target, is it a list or a single item, and how confident is each extracted value. Hundreds of questions per run, answered in one batched pass. It also runs on CPU, more slowly.

Plain code does the rest

Pagination, page decomposition and CSS selection. Four paging strategies are tried in order: a next link, a load-more button or infinite scroll, a "view all" landing page, then probing ?page= style URLs. Nothing is hard-coded per site.

  1. Plan. One LLM call turns the prompt into a target, the names sites use for such a listing, and five to nine typed fields. It is plain JSON: edit it and the next run uses your schema.
  2. Locate. The crawler proposes candidates and the LLM picks the real listing. A complete listing beats a sub-collection.
  3. Paginate. Up to the item limit you set, using whichever paging strategy the site actually has.
  4. Extract. Each page is decomposed into metadata and labelled facts, fields are mapped to elements, then every selector is graded on four sample pages and repaired until it is right.
  5. Remember. The recipe is saved per site and prompt. Repeat runs make no LLM calls. After a redesign, --revalidate re-grades the saved selectors and repairs only what broke.

What it costs, measured

This run was built from nothing on ukri.org. Every run prints its own cost by stage, and the build cost is stored in the recipe.

Cost of building a scraper compared with running it again
RunLLM callsTokens in / outLLM costTime, 100 items
Build, the first run25133,779 / 24,211$0.0697.7 min
plan, one call: $0.002 · locate, one call: $0.001 · extract, map, validate and repair, 23 calls: $0.066
Every run after, on fresh pages00$078 s
LLM cost of this run: plan 1 calls $0.0021, locate 1 calls $0.0010,
                      extract 23 calls $0.0661, total 25 calls $0.0692
LLM cost of this run: none, the recipe did all the work

Build costs vary by site. A clean site costs a few cents; a free-form blog of grants needed 46 calls, because values buried in prose need text patterns and more repair rounds. As rough arithmetic, a hundred sites like UKRI would cost about $7 to build, and nothing to run. Set LLM_PRICE_IN and LLM_PRICE_OUT to your model's prices, or to zero for a local one.

Pricing: build once for cents, then every run costs nothing.

Five targets, five sites, 100 items each

Quality here is an LLM check of ten random items per site against the page. That is a consistent proxy, not a hand-labelled ground truth.

Listing found, paging strategy, precision and recall per site
TargetSiteListing foundPagingPrecisionRecall
patentspatents.justia.com/company/linkedinnext-link100%94%
funding opportunitiesukri.org/opportunity/next-link97%91%
publicationsodi.org/en/publications/next-link96%90%
policy papersgov.uk/search/policy-papersprobe ?page=100%77%
datasetscatalog.data.govthe homepageload-more96%70%

Precision is how often a filled value is right. Recall is how many of the values shown on the page were filled. The lower recall on gov.uk and data.gov is fields left deliberately blank and marked unreliable, not wrong values: a blank is honest, a wrong value is not.

Run it yourself

Python 3.10 or newer and Node 18 or newer. An NVIDIA GPU is recommended; CPU works, more slowly. The first run downloads the laya checkpoint.

git clone https://github.com/qxZap/laya-scraper && cd laya-scraper
python -m venv .venv && .venv/Scripts/python -m pip install -r requirements.txt && npm install
cp .env.example .env     # an OpenAI or Anthropic compatible key, or a local model
.venv/Scripts/python harvest.py https://www.ukri.org/ "research funding calls" --out calls.csv

On macOS or Linux use .venv/bin/python. The LLM is needed only while a recipe is being built or re-validated, so point it at Ollama, LM Studio or vLLM and the whole thing runs on your own machine.

python harvest.py https://catalog.data.gov/ "datasets" --out datasets.json
python harvest.py https://odi.org/en/ "publications" --revalidate   # after a redesign
python scrape.py https://odi.org/en/                                # just find the listing

What it does not do

  • Weights are hand-tuned on three sites. With twenty to fifty labelled sites they could be fitted properly.
  • Some labels are noisy. The page kind is zero-shot, so an experts page can come out as "news". A confidence score is reported with each one.
  • Lists fed by a cursor-based API, with no page URLs or buttons, are out of reach. Reading the JSON API directly is the next step.
  • Recipes go stale when a site is redesigned. Re-run with --revalidate; it repairs only what broke.
  • Some fields stay blank when they vary too much across a site's pages to validate. They are marked unreliable rather than guessed.
  • Terms of service. robots.txt is respected and requests are throttled per host, but getting past a bot wall can still conflict with a site's terms. Check before crawling sites you do not own.

Point it at a site and see

The repo has the five example runs, their logs and the recipes they produced. If it fails on a site, an issue with the site, the prompt and the log is the most useful thing you can send.