How to crawl the world - do we need Google?
Pre-word
This is a rather tricky question. Or, to put it better - a question worth thinking about: how much do we actually depend on search engines?
In the age of AI we are being charged pretty much for every query or MCP call. Every web search made by an agent costs tokens, API credits, or a subscription tier. And it adds up quickly, because agents search a lot. But does it need to be this way?
Here is the thing: search engines heavily rely on robots.txt, sitemaps and feeds.
These are not proprietary APIs. They are public, old, boring conventions - and actually
we can use them on our own. The “magic index” that Google sells us access to is, to a
large degree, built from files that any of us can read for free. So the question becomes:
how far can we get without the middleman?
The free meal
Three protocols carry most of the weight:
robots.txt- the rulebook. What you may fetch, what you should not, how fast, and - this is the interesting part - where the sitemap lives.sitemap.xml- the site’s own table of contents. A full URL list, often with last-modified timestamps.- RSS / Atom feeds - the diff stream. URL, title, summary, date, author.
None of this costs anything. Nobody rate-limits you for reading a robots.txt.
And yet this is exactly the input from which the big indexes are built.
Always robots.txt first
Every time my agent touches a new domain, the first request goes to robots.txt.
Be nice to the sites which still provide the free meal - respect what they declare:
1curl -fsSL https://example.com/robots.txt
A typical file reads like this:
1User-agent: *
2Disallow: /private/
3Crawl-delay: 10
4Sitemap: https://example.com/sitemap.xml
Disallow:/Allow:- obey them for your crawler identity. This is the “please” that keeps the whole ecosystem alive.Crawl-delay: N- wait at least N seconds between requests. arxiv.org says 15, some forums say 60.Sitemap:lines - direct pointers to XML sitemaps. This is the shortcut we want.
While you are there, it costs one request to peek at /.well-known/agents.json or
/llms.txt - an emerging convention for declaring how AI agents should interact with
a site. A 404 is still the common answer, but when it exists, it is pure gold.
The sitemap - the whole URL list in one request
A sitemap gives you the full URL list without crawling a single link. The one-shot
pipeline from robots.txt to a sorted list of URLs fits in a terminal line:
1curl -fsSL https://example.com/robots.txt | grep -i sitemap: | awk '{print $2}' \
2 | while read -r s; do echo "$s"; curl -fsSL "$s"; done \
3 | grep -o '<loc>[^<]*</loc>' | sed -e 's/<loc>//g' -e 's/<\/loc>//g' | sort -u
What each stage does: pull sitemap URLs out of robots.txt, fetch each of them, strip
the <loc> tags, dedupe and sort. The grep -o '<loc>[^<]*</loc>' variant survives
minified single-line XML, which plain grep loc would mangle into one giant line.
Two gotchas worth knowing:
- Sitemap indexes. A sitemap can be a
<sitemapindex>- its<loc>entries are then child sitemaps, not pages. Expand recursively until you reach real<urlset>entries. - No
Sitemap:line? Try the common paths (/sitemap.xml,/sitemap_index.xml,/sitemap-index.xml,/wp-sitemap.xmlfor WordPress) or check for a<link rel="sitemap">in the HTML head.
And here is a thought that took me a while to appreciate: even a sitemap can be a
feed-like thing. The <lastmod> timestamps tell you what changed since your last
visit. Snapshot the URL list, diff it against the previous snapshot, and you have a
“what’s new” stream even for sites that publish no RSS at all.
RSS / Atom - URL and summary, served on a plate
Feeds are discoverable from the HTML head of the site:
1curl -fsSL https://example.com/ | grep -Eio '<link[^>]+(application/(rss|atom)\+xml)[^>]*>'
Typical tags look like this:
1<link rel="alternate" type="application/rss+xml" title="..." href="/feed.xml">
2<link rel="alternate" type="application/atom+xml" title="..." href="/atom.xml">
A feed entry gives you exactly the search-relevant triple: URL, title and summary
(plus date and author). And a nice surprise: many smaller blogs embed essentially the
whole post content in their feed entries - <content:encoded> or full Atom
<content>. Check the entry length before deciding to fetch individual pages; you
can often read every post on a site with a handful of requests instead of crawling
every URL.
The automation - teach it, don’t do it
Doing the above by hand once is fun. Doing it for every new domain is a job for a skill - a small, reusable instruction file the agent loads whenever it matches the task. Mine boils down to an order of operations:
- New domain ->
robots.txtfirst. Always. Parse rules, noteCrawl-delay, collectSitemap:lines. Be nice to the sites which still provide the free meal. - Sitemap next. Follow the pointers, expand indexes, harvest
<loc>entries. Full URL map, one or two requests. - Then feeds. Grab
<link rel="alternate" ...>from the HTML head, subscribe or fetch. That is the summary stream. - Then we search for what we want - in the sitemap by URL (URLs are surprisingly descriptive), or in the RSS feed by URL and summary - and there we go.
A few self-throttling defaults make the difference between a polite crawler and a hammer:
| Situation | Default |
|---|---|
Crawl-delay: N in robots.txt | wait >= N seconds between requests |
| No crawl-delay, plain pages | >= 30 s between requests to the same host |
| Cloudflare / anti-bot wall | hard-block that host for >= 5 minutes, no retries |
| Sitemaps / feeds | 1-2 requests replace dozens of page fetches - prefer |
Rule of thumb: minimize request count first, throttle what remains. If your plan needs more than a handful of requests per host, you probably missed a sitemap or a feed.
One more warning: sometimes squeezing lemon is more overkill
than it is really worth. The temptation is always there - the SPA shell is empty,
robots.txt disallows the interesting path, and “just this once” you could spoof a
browser user-agent, rotate a few proxies, bolt stealth plugins onto headless chromium.
Before you start crawling and ignoring robots.txt, make sure it is worth risking the
anti-crawling techniques: TLS and canvas fingerprinting, request-cadence detection,
Cloudflare/DataDome/Akamai walls - and those hit the VPS your crawler runs on far
more often than your household: expect the ban plus an angry ticket from your
hosting provider - and
your IP or domain ending up on a Spamhaus-style blocklist, which hurts far beyond one
scraping session once your mail starts bouncing :-) All of that, usually
to obtain content that the same site offers through a sitemap or a feed you were
invited to read. The polite path covers the bulk of the value for a fraction of the
effort - and when the content truly matters and is locked away, the honest moves are
the site’s own API, asking the author, or paying for the data. There is also a
commons argument: every agent that misbehaves today gives sites a reason to wall
themselves off tomorrow, and the free meal shrinks for everyone who comes after.
Keeping it - the local index
Once the URLs and summaries arrive, where do they live? Two honest options:
ChromaDB - basic and easy. Good enough for a personal corpus, semantic search included. Start the server as a container - single process, named volume for the data, port 8000 for the HTTP client:
1podman container run --name chromadb \
2 -p 8000:8000 \
3 -v chroma_data:/data \
4 docker.io/chromadb/chroma
The sync loop for a feed is a few dozen lines of Python: parse the RSS, upsert by entry id (idempotent - reruns fix themselves), remember a watermark timestamp in collection metadata so the next sync fetches only the delta:
1import re, time, urllib.request
2import xml.etree.ElementTree as ET
3import chromadb
4
5col = chromadb.HttpClient(host="127.0.0.1", port=8000) \
6 .get_or_create_collection("feeds", metadata={"last_sync": 0})
7last = int((col.metadata or {}).get("last_sync", 0))
8
9with urllib.request.urlopen(FEED_URL, timeout=60) as r:
10 items = ET.fromstring(r.read()).findall(".//item")
11
12ids, docs, metas = [], [], []
13for it in items:
14 guid = (it.findtext("guid") or "").strip()
15 title = (it.findtext("title") or "").strip()
16 desc = re.sub(r"\s+", " ", re.sub(r"<[^>]+>", " ",
17 it.findtext("description") or "")).strip()
18 ids.append(guid)
19 docs.append(f"{title}\n\n{desc}".strip())
20 metas.append({"link": (it.findtext("link") or "").strip()})
21
22if ids:
23 col.upsert(ids=ids, documents=docs, metadatas=metas)
24col.modify(metadata={"last_sync": int(time.time())})
One caveat: ChromaDB does not really work when you query big data - feed entries
fit in one record, but a whole crawled page would end up as one giant embedding
that drowns the details. So chunk bigger records: split the text into blocks of
1024 or 4096 characters, upsert one record per block, id like
f"{url}#{offset}", URL in the metadata - and every hit points straight back
to its source.
Searching it later is a one-liner:
1r = col.query(query_texts=["hugo template lookup order"], n_results=10)
2for iid, m, d in zip(r["ids"][0], r["metadatas"][0], r["distances"][0]):
3 print(f"{d:.3f} | {m['link']}")
Add structured filters on the metadata you stored - where={"published": {"$gte": since}}
to narrow by date, or where={"tags": "hugo"} if you tagged each page - instead of
trusting the vectors alone.
OpenSearch (or Elasticsearch, or a similar Lucene-based engine) - when you outgrow the simple case. Real BM25 scoring, filters, aggregations, dashboards. It runs fine as a single-node container:
1podman container run --name opensearch-db \
2 -e "discovery.type=single-node" \
3 -e OPENSEARCH_INITIAL_ADMIN_PASSWORD=SECRET_HERE \
4 -p 9200:9200 -p 9300:9300 \
5 -v opensearch_data:/usr/share/opensearch \
6 docker.io/opensearchproject/opensearch
Index the harvest as records too - page content in blocks of 1024 or 4096
characters, the source URL and any tags as plain fields - and search full text
over content, no vectors involved. More ceremony than the ChromaDB one-liner:
the official client, plus the dev container’s self-signed TLS:
1from opensearchpy import OpenSearch
2
3client = OpenSearch(
4 hosts=[{"host": "127.0.0.1", "port": 9200}],
5 http_auth=("admin", "SECRET_HERE"),
6 use_ssl=True, verify_certs=False,
7)
8
9r = client.search(index="pages", body={
10 "size": 10,
11 "query": {"bool": {
12 "must": {"match": {"content": "hugo template lookup order"}},
13 "filter": {"term": {"tags": "hugo"}},
14 }},
15})
16for hit in r["hits"]["hits"]:
17 print(f'{hit["_score"]:.3f} | {hit["_source"]["link"]}')
Same shape as the ChromaDB loop, with the BM25 _score in place of similarity
distance. Drop the filter clause for plain full text; the term query on tags
only behaves when the field is mapped as keyword - not analyzed into separate
words. And when bool is not enough, the whole query DSL lives in the body:
match_phrase for word order, range for dates, aggregations for counts.
For “I want to find that blog post I half-remember”, ChromaDB is honestly enough.
Reach for OpenSearch when you want query languages, exact-term matching at scale,
or you are indexing millions of documents. There is also the middle ground I use
daily: ripgrep over plain markdown files when the corpus is just files on disk.
What is still missing - the honest part
Let me not oversell this. What the DIY pipeline does not give you:
- No global relevance ranking. Google’s index spans the whole web and ranks it. My pipeline answers “what does this site have” - you need to know the domain (or a seed page) first. Discovery across the whole web is a different, harder game.
- Stale or missing data. Sitemaps lie (or rather, lag). Not every site has feeds. Some sitemaps are generated once a year and never touched again.
- JavaScript-only sites. SPA shells return an empty
<div id="root">to curl. Headless chromium in dump-DOM mode (--dump-dom --virtual-time-budget=10000) renders them, at the cost of a heavy page load - use sparingly, and throttle. - You are the ranking function. The agent reading ten titles and summaries and picking the right one is, effectively, doing the ranking. That works remarkably well in practice - but it is a different trade-off than a tuned index.
What is left?
The pipeline above answers “where on this site is X”. A few problems it deliberately leaves open - and where I think the interesting work is:
Rankings
What we do not have is PageRank. Google’s relevance is built on the link graph - who points at whom, weighted and iterated - plus decades of click-through and quality signals. Locally we have none of that. But do we actually need it? A per-site corpus is small; a hundred titles fit in an agent’s context window without any ranking at all. Where it starts to matter is choosing between sources, and there are cheap proxies that go a long way:
- recency - feeds give it for free,
- your own hit rate - did this source actually answer anything last time? Keep the score per domain and you get a personal relevance signal no crawler can fake,
- curated seeds - the sites you bookmarked once and keep coming back to.
And when none of that settles it, the agent reads ten titles and picks. That works remarkably well - but every such judgment is an AI call, so caching the verdicts (“this blog is a good source for X”) is where the real savings live.
Crawl only what is worth crawling
We do not need all the world’s slop. A crawl budget - even a generous one - should follow worthiness, not completeness. The signals are the same ones as above, applied before spending any request:
- does the site bother to publish a sitemap and feeds? That alone is a statement of intent: it wants to be read by machines. Sites without any of it force you into page-by-page guessing - the most expensive kind of crawl - and often are not worth the guess,
- past usefulness - a domain whose pages answered your questions five times earns more budget than one that answered zero,
- who references it - your own notes, bookmarks and previous answers are a link graph, just a private one.
Update the per-domain score every time a fetch proves useful or useless and you end up with a tiny, self-correcting priority list. Effectively a PageRank where the “links” are your own queries that led somewhere.
Recrawl policies
The last open piece: when to come back. Feeds are almost self-solving - poll them,
use conditional requests (ETag, If-Modified-Since), and an unchanged site costs
you a 304. Sitemaps are trickier: how often to re-fetch depends on the observed
change rate, which you get for free from the <lastmod> diffs mentioned above - a
news site and a personal blog clearly deserve different schedules. There is also the
question of expiring stale entries from the local index versus keeping them as
“was true once” history. Each of these deserves more space than a bullet - maybe
the subject of the next post.
So - do we need Google?
For finding anything on the whole web - probably yes, for now. The long tail of discovery is what the big indexes are genuinely good at.
But for the question I actually ask dozens of times a day - “where on this site is
the thing I remember?” - the answer is: we do not need a search engine billing us
per query. robots.txt, a sitemap and a feed are a site’s own index, published for
free, machine-readable, waiting to be used. A skill that reads them in order and a
local ChromaDB to keep the harvest replace a surprisingly large chunk of what we
have been paying for. And unlike a commercial API, the free meal comes with an
etiquette: read robots.txt, honor Crawl-delay, prefer sitemaps and feeds over
page-by-page crawling.
One last thing, since we are all someone else’s crawler target: publish your own
sitemap and feeds. This very site does - Hugo generates /sitemap.xml and
/index.xml (RSS) and /atom.xml out of the box, zero extra work. Do it, and the
agents of the world can find your content without paying a middleman for the
privilege.