How to crawl the world - do we need Google?

Pre-word

This is a rather tricky question. Or, to put it better - a question worth thinking about: how much do we actually depend on search engines?

In the age of AI we are being charged pretty much for every query or MCP call. Every web search made by an agent costs tokens, API credits, or a subscription tier. And it adds up quickly, because agents search a lot. But does it need to be this way?

Here is the thing: search engines heavily rely on robots.txt, sitemaps and feeds. These are not proprietary APIs. They are public, old, boring conventions - and actually we can use them on our own. The “magic index” that Google sells us access to is, to a large degree, built from files that any of us can read for free. So the question becomes: how far can we get without the middleman?

The free meal

Three protocols carry most of the weight:

  1. robots.txt - the rulebook. What you may fetch, what you should not, how fast, and - this is the interesting part - where the sitemap lives.
  2. sitemap.xml - the site’s own table of contents. A full URL list, often with last-modified timestamps.
  3. RSS / Atom feeds - the diff stream. URL, title, summary, date, author.

None of this costs anything. Nobody rate-limits you for reading a robots.txt. And yet this is exactly the input from which the big indexes are built.

Always robots.txt first

Every time my agent touches a new domain, the first request goes to robots.txt. Be nice to the sites which still provide the free meal - respect what they declare:

1curl -fsSL https://example.com/robots.txt

A typical file reads like this:

1User-agent: *
2Disallow: /private/
3Crawl-delay: 10
4Sitemap: https://example.com/sitemap.xml

While you are there, it costs one request to peek at /.well-known/agents.json or /llms.txt - an emerging convention for declaring how AI agents should interact with a site. A 404 is still the common answer, but when it exists, it is pure gold.

The sitemap - the whole URL list in one request

A sitemap gives you the full URL list without crawling a single link. The one-shot pipeline from robots.txt to a sorted list of URLs fits in a terminal line:

1curl -fsSL https://example.com/robots.txt | grep -i sitemap: | awk '{print $2}' \
2  | while read -r s; do echo "$s"; curl -fsSL "$s"; done \
3  | grep -o '<loc>[^<]*</loc>' | sed -e 's/<loc>//g' -e 's/<\/loc>//g' | sort -u

What each stage does: pull sitemap URLs out of robots.txt, fetch each of them, strip the <loc> tags, dedupe and sort. The grep -o '<loc>[^<]*</loc>' variant survives minified single-line XML, which plain grep loc would mangle into one giant line.

Two gotchas worth knowing:

And here is a thought that took me a while to appreciate: even a sitemap can be a feed-like thing. The <lastmod> timestamps tell you what changed since your last visit. Snapshot the URL list, diff it against the previous snapshot, and you have a “what’s new” stream even for sites that publish no RSS at all.

RSS / Atom - URL and summary, served on a plate

Feeds are discoverable from the HTML head of the site:

1curl -fsSL https://example.com/ | grep -Eio '<link[^>]+(application/(rss|atom)\+xml)[^>]*>'

Typical tags look like this:

1<link rel="alternate" type="application/rss+xml" title="..." href="/feed.xml">
2<link rel="alternate" type="application/atom+xml" title="..." href="/atom.xml">

A feed entry gives you exactly the search-relevant triple: URL, title and summary (plus date and author). And a nice surprise: many smaller blogs embed essentially the whole post content in their feed entries - <content:encoded> or full Atom <content>. Check the entry length before deciding to fetch individual pages; you can often read every post on a site with a handful of requests instead of crawling every URL.

The automation - teach it, don’t do it

Doing the above by hand once is fun. Doing it for every new domain is a job for a skill - a small, reusable instruction file the agent loads whenever it matches the task. Mine boils down to an order of operations:

  1. New domain -> robots.txt first. Always. Parse rules, note Crawl-delay, collect Sitemap: lines. Be nice to the sites which still provide the free meal.
  2. Sitemap next. Follow the pointers, expand indexes, harvest <loc> entries. Full URL map, one or two requests.
  3. Then feeds. Grab <link rel="alternate" ...> from the HTML head, subscribe or fetch. That is the summary stream.
  4. Then we search for what we want - in the sitemap by URL (URLs are surprisingly descriptive), or in the RSS feed by URL and summary - and there we go.

A few self-throttling defaults make the difference between a polite crawler and a hammer:

SituationDefault
Crawl-delay: N in robots.txtwait >= N seconds between requests
No crawl-delay, plain pages>= 30 s between requests to the same host
Cloudflare / anti-bot wallhard-block that host for >= 5 minutes, no retries
Sitemaps / feeds1-2 requests replace dozens of page fetches - prefer

Rule of thumb: minimize request count first, throttle what remains. If your plan needs more than a handful of requests per host, you probably missed a sitemap or a feed.

One more warning: sometimes squeezing lemon is more overkill than it is really worth. The temptation is always there - the SPA shell is empty, robots.txt disallows the interesting path, and “just this once” you could spoof a browser user-agent, rotate a few proxies, bolt stealth plugins onto headless chromium. Before you start crawling and ignoring robots.txt, make sure it is worth risking the anti-crawling techniques: TLS and canvas fingerprinting, request-cadence detection, Cloudflare/DataDome/Akamai walls - and those hit the VPS your crawler runs on far more often than your household: expect the ban plus an angry ticket from your hosting provider - and your IP or domain ending up on a Spamhaus-style blocklist, which hurts far beyond one scraping session once your mail starts bouncing :-) All of that, usually to obtain content that the same site offers through a sitemap or a feed you were invited to read. The polite path covers the bulk of the value for a fraction of the effort - and when the content truly matters and is locked away, the honest moves are the site’s own API, asking the author, or paying for the data. There is also a commons argument: every agent that misbehaves today gives sites a reason to wall themselves off tomorrow, and the free meal shrinks for everyone who comes after.

Keeping it - the local index

Once the URLs and summaries arrive, where do they live? Two honest options:

ChromaDB - basic and easy. Good enough for a personal corpus, semantic search included. Start the server as a container - single process, named volume for the data, port 8000 for the HTTP client:

1podman container run --name chromadb \
2    -p 8000:8000 \
3    -v chroma_data:/data \
4    docker.io/chromadb/chroma

The sync loop for a feed is a few dozen lines of Python: parse the RSS, upsert by entry id (idempotent - reruns fix themselves), remember a watermark timestamp in collection metadata so the next sync fetches only the delta:

 1import re, time, urllib.request
 2import xml.etree.ElementTree as ET
 3import chromadb
 4
 5col = chromadb.HttpClient(host="127.0.0.1", port=8000) \
 6    .get_or_create_collection("feeds", metadata={"last_sync": 0})
 7last = int((col.metadata or {}).get("last_sync", 0))
 8
 9with urllib.request.urlopen(FEED_URL, timeout=60) as r:
10    items = ET.fromstring(r.read()).findall(".//item")
11
12ids, docs, metas = [], [], []
13for it in items:
14    guid = (it.findtext("guid") or "").strip()
15    title = (it.findtext("title") or "").strip()
16    desc = re.sub(r"\s+", " ", re.sub(r"<[^>]+>", " ",
17                  it.findtext("description") or "")).strip()
18    ids.append(guid)
19    docs.append(f"{title}\n\n{desc}".strip())
20    metas.append({"link": (it.findtext("link") or "").strip()})
21
22if ids:
23    col.upsert(ids=ids, documents=docs, metadatas=metas)
24col.modify(metadata={"last_sync": int(time.time())})

One caveat: ChromaDB does not really work when you query big data - feed entries fit in one record, but a whole crawled page would end up as one giant embedding that drowns the details. So chunk bigger records: split the text into blocks of 1024 or 4096 characters, upsert one record per block, id like f"{url}#{offset}", URL in the metadata - and every hit points straight back to its source.

Searching it later is a one-liner:

1r = col.query(query_texts=["hugo template lookup order"], n_results=10)
2for iid, m, d in zip(r["ids"][0], r["metadatas"][0], r["distances"][0]):
3    print(f"{d:.3f} | {m['link']}")

Add structured filters on the metadata you stored - where={"published": {"$gte": since}} to narrow by date, or where={"tags": "hugo"} if you tagged each page - instead of trusting the vectors alone.

OpenSearch (or Elasticsearch, or a similar Lucene-based engine) - when you outgrow the simple case. Real BM25 scoring, filters, aggregations, dashboards. It runs fine as a single-node container:

1podman container run --name opensearch-db \
2    -e "discovery.type=single-node" \
3    -e OPENSEARCH_INITIAL_ADMIN_PASSWORD=SECRET_HERE \
4    -p 9200:9200 -p 9300:9300 \
5    -v opensearch_data:/usr/share/opensearch \
6    docker.io/opensearchproject/opensearch

Index the harvest as records too - page content in blocks of 1024 or 4096 characters, the source URL and any tags as plain fields - and search full text over content, no vectors involved. More ceremony than the ChromaDB one-liner: the official client, plus the dev container’s self-signed TLS:

 1from opensearchpy import OpenSearch
 2
 3client = OpenSearch(
 4    hosts=[{"host": "127.0.0.1", "port": 9200}],
 5    http_auth=("admin", "SECRET_HERE"),
 6    use_ssl=True, verify_certs=False,
 7)
 8
 9r = client.search(index="pages", body={
10    "size": 10,
11    "query": {"bool": {
12        "must": {"match": {"content": "hugo template lookup order"}},
13        "filter": {"term": {"tags": "hugo"}},
14    }},
15})
16for hit in r["hits"]["hits"]:
17    print(f'{hit["_score"]:.3f} | {hit["_source"]["link"]}')

Same shape as the ChromaDB loop, with the BM25 _score in place of similarity distance. Drop the filter clause for plain full text; the term query on tags only behaves when the field is mapped as keyword - not analyzed into separate words. And when bool is not enough, the whole query DSL lives in the body: match_phrase for word order, range for dates, aggregations for counts.

For “I want to find that blog post I half-remember”, ChromaDB is honestly enough. Reach for OpenSearch when you want query languages, exact-term matching at scale, or you are indexing millions of documents. There is also the middle ground I use daily: ripgrep over plain markdown files when the corpus is just files on disk.

What is still missing - the honest part

Let me not oversell this. What the DIY pipeline does not give you:

What is left?

The pipeline above answers “where on this site is X”. A few problems it deliberately leaves open - and where I think the interesting work is:

Rankings

What we do not have is PageRank. Google’s relevance is built on the link graph - who points at whom, weighted and iterated - plus decades of click-through and quality signals. Locally we have none of that. But do we actually need it? A per-site corpus is small; a hundred titles fit in an agent’s context window without any ranking at all. Where it starts to matter is choosing between sources, and there are cheap proxies that go a long way:

And when none of that settles it, the agent reads ten titles and picks. That works remarkably well - but every such judgment is an AI call, so caching the verdicts (“this blog is a good source for X”) is where the real savings live.

Crawl only what is worth crawling

We do not need all the world’s slop. A crawl budget - even a generous one - should follow worthiness, not completeness. The signals are the same ones as above, applied before spending any request:

Update the per-domain score every time a fetch proves useful or useless and you end up with a tiny, self-correcting priority list. Effectively a PageRank where the “links” are your own queries that led somewhere.

Recrawl policies

The last open piece: when to come back. Feeds are almost self-solving - poll them, use conditional requests (ETag, If-Modified-Since), and an unchanged site costs you a 304. Sitemaps are trickier: how often to re-fetch depends on the observed change rate, which you get for free from the <lastmod> diffs mentioned above - a news site and a personal blog clearly deserve different schedules. There is also the question of expiring stale entries from the local index versus keeping them as “was true once” history. Each of these deserves more space than a bullet - maybe the subject of the next post.

So - do we need Google?

For finding anything on the whole web - probably yes, for now. The long tail of discovery is what the big indexes are genuinely good at.

But for the question I actually ask dozens of times a day - “where on this site is the thing I remember?” - the answer is: we do not need a search engine billing us per query. robots.txt, a sitemap and a feed are a site’s own index, published for free, machine-readable, waiting to be used. A skill that reads them in order and a local ChromaDB to keep the harvest replace a surprisingly large chunk of what we have been paying for. And unlike a commercial API, the free meal comes with an etiquette: read robots.txt, honor Crawl-delay, prefer sitemaps and feeds over page-by-page crawling.

One last thing, since we are all someone else’s crawler target: publish your own sitemap and feeds. This very site does - Hugo generates /sitemap.xml and /index.xml (RSS) and /atom.xml out of the box, zero extra work. Do it, and the agents of the world can find your content without paying a middleman for the privilege.

Tags: #Crawling #Search #Ai #Automation