<?xml version="1.0" encoding="utf-8" standalone="yes"?><feed xmlns="http://www.w3.org/2005/Atom"><title>Automation</title><id>https://polgrabia.me/tags/automation/</id><updated>2026-10-04T19:00:00+02:00</updated><link rel="self" href="https://polgrabia.me/tags/automation/"/><link rel="alternate" type="text/html" href="https://polgrabia.me/tags/automation/"/><author><name>Tomasz Półgrabia</name></author><generator>Hugo</generator><entry><title>How to crawl the world - do we need Google?</title><id>https://polgrabia.me/posts/20261004-how-to-crawl-the-world-do-we-need-google/</id><link rel="alternate" href="https://polgrabia.me/posts/20261004-how-to-crawl-the-world-do-we-need-google/"/><published>2026-10-04T19:00:00+02:00</published><updated>2026-10-04T19:00:00+02:00</updated><category term="crawling"/><category term="search"/><category term="ai"/><category term="automation"/><content type="html">&lt;h1 id="pre-word"&gt;Pre-word&lt;/h1&gt;
&lt;p&gt;This is a rather tricky question. Or, to put it better - a question worth thinking about:
how much do we actually depend on search engines?&lt;/p&gt;
&lt;p&gt;In the age of AI we are being charged pretty much for every query or MCP call. Every web
search made by an agent costs tokens, API credits, or a subscription tier. And it adds up
quickly, because agents search a &lt;em&gt;lot&lt;/em&gt;. But does it need to be this way?&lt;/p&gt;
&lt;p&gt;Here is the thing: search engines heavily rely on &lt;code&gt;robots.txt&lt;/code&gt;, sitemaps and feeds.
These are not proprietary APIs. They are public, old, boring conventions - and actually
we can use them on our own. The &amp;ldquo;magic index&amp;rdquo; that Google sells us access to is, to a
large degree, built from files that any of us can read for free. So the question becomes:
how far can we get without the middleman?&lt;/p&gt;
&lt;h1 id="the-free-meal"&gt;The free meal&lt;/h1&gt;
&lt;p&gt;Three protocols carry most of the weight:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;code&gt;robots.txt&lt;/code&gt; - the rulebook. What you may fetch, what you should not, how fast,
and - this is the interesting part - &lt;em&gt;where the sitemap lives&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;sitemap.xml&lt;/code&gt; - the site&amp;rsquo;s own table of contents. A full URL list, often with
last-modified timestamps.&lt;/li&gt;
&lt;li&gt;RSS / Atom feeds - the diff stream. URL, title, summary, date, author.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;None of this costs anything. Nobody rate-limits you for reading a &lt;code&gt;robots.txt&lt;/code&gt;.
And yet this is exactly the input from which the big indexes are built.&lt;/p&gt;
&lt;h1 id="always-robotstxt-first"&gt;Always robots.txt first&lt;/h1&gt;
&lt;p&gt;Every time my agent touches a new domain, the first request goes to &lt;code&gt;robots.txt&lt;/code&gt;.
Be nice to the sites which still provide the free meal - respect what they declare:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;1&lt;/span&gt;&lt;span&gt;curl -fsSL https://example.com/robots.txt
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A typical file reads like this:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;1&lt;/span&gt;&lt;span&gt;User-agent: *
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;2&lt;/span&gt;&lt;span&gt;Disallow: /private/
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;3&lt;/span&gt;&lt;span&gt;Crawl-delay: 10
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;4&lt;/span&gt;&lt;span&gt;Sitemap: https://example.com/sitemap.xml
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;ul&gt;
&lt;li&gt;&lt;code&gt;Disallow:&lt;/code&gt; / &lt;code&gt;Allow:&lt;/code&gt; - obey them for your crawler identity. This is the &amp;ldquo;please&amp;rdquo;
that keeps the whole ecosystem alive.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Crawl-delay: N&lt;/code&gt; - wait at least N seconds between requests. arxiv.org says 15,
some forums say 60.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Sitemap:&lt;/code&gt; lines - direct pointers to XML sitemaps. This is the shortcut we want.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;While you are there, it costs one request to peek at &lt;code&gt;/.well-known/agents.json&lt;/code&gt; or
&lt;code&gt;/llms.txt&lt;/code&gt; - an emerging convention for declaring how AI agents should interact with
a site. A 404 is still the common answer, but when it exists, it is pure gold.&lt;/p&gt;
&lt;h1 id="the-sitemap---the-whole-url-list-in-one-request"&gt;The sitemap - the whole URL list in one request&lt;/h1&gt;
&lt;p&gt;A sitemap gives you the full URL list &lt;em&gt;without crawling a single link&lt;/em&gt;. The one-shot
pipeline from &lt;code&gt;robots.txt&lt;/code&gt; to a sorted list of URLs fits in a terminal line:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;1&lt;/span&gt;&lt;span&gt;curl -fsSL https://example.com/robots.txt | grep -i sitemap: | awk &lt;span style="color:#e6db74"&gt;&amp;#39;{print $2}&amp;#39;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;2&lt;/span&gt;&lt;span&gt; | &lt;span style="color:#66d9ef"&gt;while&lt;/span&gt; read -r s; &lt;span style="color:#66d9ef"&gt;do&lt;/span&gt; echo &lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;$s&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;; curl -fsSL &lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;$s&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;; &lt;span style="color:#66d9ef"&gt;done&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;3&lt;/span&gt;&lt;span&gt; | grep -o &lt;span style="color:#e6db74"&gt;&amp;#39;&amp;lt;loc&amp;gt;[^&amp;lt;]*&amp;lt;/loc&amp;gt;&amp;#39;&lt;/span&gt; | sed -e &lt;span style="color:#e6db74"&gt;&amp;#39;s/&amp;lt;loc&amp;gt;//g&amp;#39;&lt;/span&gt; -e &lt;span style="color:#e6db74"&gt;&amp;#39;s/&amp;lt;\/loc&amp;gt;//g&amp;#39;&lt;/span&gt; | sort -u
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What each stage does: pull sitemap URLs out of &lt;code&gt;robots.txt&lt;/code&gt;, fetch each of them, strip
the &lt;code&gt;&amp;lt;loc&amp;gt;&lt;/code&gt; tags, dedupe and sort. The &lt;code&gt;grep -o '&amp;lt;loc&amp;gt;[^&amp;lt;]*&amp;lt;/loc&amp;gt;'&lt;/code&gt; variant survives
minified single-line XML, which plain &lt;code&gt;grep loc&lt;/code&gt; would mangle into one giant line.&lt;/p&gt;
&lt;p&gt;Two gotchas worth knowing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Sitemap indexes.&lt;/strong&gt; A sitemap can be a &lt;code&gt;&amp;lt;sitemapindex&amp;gt;&lt;/code&gt; - its &lt;code&gt;&amp;lt;loc&amp;gt;&lt;/code&gt; entries are
then &lt;em&gt;child sitemaps&lt;/em&gt;, not pages. Expand recursively until you reach real
&lt;code&gt;&amp;lt;urlset&amp;gt;&lt;/code&gt; entries.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No &lt;code&gt;Sitemap:&lt;/code&gt; line?&lt;/strong&gt; Try the common paths (&lt;code&gt;/sitemap.xml&lt;/code&gt;, &lt;code&gt;/sitemap_index.xml&lt;/code&gt;,
&lt;code&gt;/sitemap-index.xml&lt;/code&gt;, &lt;code&gt;/wp-sitemap.xml&lt;/code&gt; for WordPress) or check for a
&lt;code&gt;&amp;lt;link rel=&amp;quot;sitemap&amp;quot;&amp;gt;&lt;/code&gt; in the HTML head.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And here is a thought that took me a while to appreciate: &lt;strong&gt;even a sitemap can be a
feed-like thing&lt;/strong&gt;. The &lt;code&gt;&amp;lt;lastmod&amp;gt;&lt;/code&gt; timestamps tell you what changed since your last
visit. Snapshot the URL list, diff it against the previous snapshot, and you have a
&amp;ldquo;what&amp;rsquo;s new&amp;rdquo; stream even for sites that publish no RSS at all.&lt;/p&gt;
&lt;h1 id="rss--atom---url-and-summary-served-on-a-plate"&gt;RSS / Atom - URL and summary, served on a plate&lt;/h1&gt;
&lt;p&gt;Feeds are discoverable from the HTML head of the site:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;1&lt;/span&gt;&lt;span&gt;curl -fsSL https://example.com/ | grep -Eio &lt;span style="color:#e6db74"&gt;&amp;#39;&amp;lt;link[^&amp;gt;]+(application/(rss|atom)\+xml)[^&amp;gt;]*&amp;gt;&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Typical tags look like this:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-html" data-lang="html"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;1&lt;/span&gt;&lt;span&gt;&amp;lt;&lt;span style="color:#f92672"&gt;link&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;rel&lt;/span&gt;&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;type&lt;/span&gt;&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;application/rss+xml&amp;#34;&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;title&lt;/span&gt;&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;...&amp;#34;&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;href&lt;/span&gt;&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;/feed.xml&amp;#34;&lt;/span&gt;&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;2&lt;/span&gt;&lt;span&gt;&amp;lt;&lt;span style="color:#f92672"&gt;link&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;rel&lt;/span&gt;&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;type&lt;/span&gt;&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;application/atom+xml&amp;#34;&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;title&lt;/span&gt;&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;...&amp;#34;&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;href&lt;/span&gt;&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;/atom.xml&amp;#34;&lt;/span&gt;&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A feed entry gives you exactly the search-relevant triple: &lt;strong&gt;URL, title and summary&lt;/strong&gt;
(plus date and author). And a nice surprise: many smaller blogs embed essentially the
&lt;em&gt;whole post content&lt;/em&gt; in their feed entries - &lt;code&gt;&amp;lt;content:encoded&amp;gt;&lt;/code&gt; or full Atom
&lt;code&gt;&amp;lt;content&amp;gt;&lt;/code&gt;. Check the entry length before deciding to fetch individual pages; you
can often read every post on a site with a handful of requests instead of crawling
every URL.&lt;/p&gt;
&lt;h1 id="the-automation---teach-it-dont-do-it"&gt;The automation - teach it, don&amp;rsquo;t do it&lt;/h1&gt;
&lt;p&gt;Doing the above by hand once is fun. Doing it for every new domain is a job for a
&lt;em&gt;skill&lt;/em&gt; - a small, reusable instruction file the agent loads whenever it matches the
task. Mine boils down to an order of operations:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;New domain -&amp;gt; &lt;code&gt;robots.txt&lt;/code&gt; first.&lt;/strong&gt; Always. Parse rules, note &lt;code&gt;Crawl-delay&lt;/code&gt;,
collect &lt;code&gt;Sitemap:&lt;/code&gt; lines. Be nice to the sites which still provide the free meal.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sitemap next.&lt;/strong&gt; Follow the pointers, expand indexes, harvest &lt;code&gt;&amp;lt;loc&amp;gt;&lt;/code&gt; entries.
Full URL map, one or two requests.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Then feeds.&lt;/strong&gt; Grab &lt;code&gt;&amp;lt;link rel=&amp;quot;alternate&amp;quot; ...&amp;gt;&lt;/code&gt; from the HTML head, subscribe
or fetch. That is the summary stream.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Then we search for what we want&lt;/strong&gt; - in the sitemap &lt;em&gt;by URL&lt;/em&gt; (URLs are
surprisingly descriptive), or in the RSS feed &lt;em&gt;by URL and summary&lt;/em&gt; - and there
we go.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A few self-throttling defaults make the difference between a polite crawler and a
hammer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Crawl-delay: N&lt;/code&gt; in robots.txt&lt;/td&gt;
&lt;td&gt;wait &amp;gt;= N seconds between requests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No crawl-delay, plain pages&lt;/td&gt;
&lt;td&gt;&amp;gt;= 30 s between requests to the same host&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloudflare / anti-bot wall&lt;/td&gt;
&lt;td&gt;hard-block that host for &amp;gt;= 5 minutes, no retries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sitemaps / feeds&lt;/td&gt;
&lt;td&gt;1-2 requests replace dozens of page fetches - prefer&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Rule of thumb: &lt;strong&gt;minimize request count first, throttle what remains.&lt;/strong&gt; If your plan
needs more than a handful of requests per host, you probably missed a sitemap or a feed.&lt;/p&gt;
&lt;p&gt;One more warning: sometimes squeezing lemon is more overkill
than it is really worth. The temptation is always there - the SPA shell is empty,
&lt;code&gt;robots.txt&lt;/code&gt; disallows the interesting path, and &amp;ldquo;just this once&amp;rdquo; you could spoof a
browser user-agent, rotate a few proxies, bolt stealth plugins onto headless chromium.
Before you start crawling and ignoring &lt;code&gt;robots.txt&lt;/code&gt;, make sure it is worth risking the
anti-crawling techniques: TLS and canvas fingerprinting, request-cadence detection,
Cloudflare/DataDome/Akamai walls - and those hit the VPS your crawler runs on far
more often than your household: expect the ban plus an angry ticket from your
hosting provider - and
your IP or domain ending up on a Spamhaus-style blocklist, which hurts far beyond one
scraping session once your mail starts bouncing :-) All of that, usually
to obtain content that the same site offers through a sitemap or a feed you were
&lt;em&gt;invited&lt;/em&gt; to read. The polite path covers the bulk of the value for a fraction of the
effort - and when the content truly matters and is locked away, the honest moves are
the site&amp;rsquo;s own API, asking the author, or paying for the data. There is also a
commons argument: every agent that misbehaves today gives sites a reason to wall
themselves off tomorrow, and the free meal shrinks for everyone who comes after.&lt;/p&gt;
&lt;h1 id="keeping-it---the-local-index"&gt;Keeping it - the local index&lt;/h1&gt;
&lt;p&gt;Once the URLs and summaries arrive, where do they live? Two honest options:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;ChromaDB&lt;/strong&gt; - basic and easy. Good enough for a personal corpus, semantic search
included. Start the server as a container - single process, named volume for the
data, port 8000 for the HTTP client:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;1&lt;/span&gt;&lt;span&gt;podman container run --name chromadb &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;2&lt;/span&gt;&lt;span&gt; -p 8000:8000 &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;3&lt;/span&gt;&lt;span&gt; -v chroma_data:/data &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;4&lt;/span&gt;&lt;span&gt; docker.io/chromadb/chroma
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The sync loop for a feed is a few dozen lines of Python: parse the RSS,
upsert by entry id (idempotent - reruns fix themselves), remember a watermark
timestamp in collection metadata so the next sync fetches only the delta:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 1&lt;/span&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; re&lt;span style="color:#f92672"&gt;,&lt;/span&gt; time&lt;span style="color:#f92672"&gt;,&lt;/span&gt; urllib.request
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 2&lt;/span&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; xml.etree.ElementTree &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; ET
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 3&lt;/span&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; chromadb
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 4&lt;/span&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 5&lt;/span&gt;&lt;span&gt;col &lt;span style="color:#f92672"&gt;=&lt;/span&gt; chromadb&lt;span style="color:#f92672"&gt;.&lt;/span&gt;HttpClient(host&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;127.0.0.1&amp;#34;&lt;/span&gt;, port&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;8000&lt;/span&gt;) \
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 6&lt;/span&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;.&lt;/span&gt;get_or_create_collection(&lt;span style="color:#e6db74"&gt;&amp;#34;feeds&amp;#34;&lt;/span&gt;, metadata&lt;span style="color:#f92672"&gt;=&lt;/span&gt;{&lt;span style="color:#e6db74"&gt;&amp;#34;last_sync&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;})
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 7&lt;/span&gt;&lt;span&gt;last &lt;span style="color:#f92672"&gt;=&lt;/span&gt; int((col&lt;span style="color:#f92672"&gt;.&lt;/span&gt;metadata &lt;span style="color:#f92672"&gt;or&lt;/span&gt; {})&lt;span style="color:#f92672"&gt;.&lt;/span&gt;get(&lt;span style="color:#e6db74"&gt;&amp;#34;last_sync&amp;#34;&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;))
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 8&lt;/span&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 9&lt;/span&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;with&lt;/span&gt; urllib&lt;span style="color:#f92672"&gt;.&lt;/span&gt;request&lt;span style="color:#f92672"&gt;.&lt;/span&gt;urlopen(FEED_URL, timeout&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;60&lt;/span&gt;) &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; r:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;10&lt;/span&gt;&lt;span&gt; items &lt;span style="color:#f92672"&gt;=&lt;/span&gt; ET&lt;span style="color:#f92672"&gt;.&lt;/span&gt;fromstring(r&lt;span style="color:#f92672"&gt;.&lt;/span&gt;read())&lt;span style="color:#f92672"&gt;.&lt;/span&gt;findall(&lt;span style="color:#e6db74"&gt;&amp;#34;.//item&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;11&lt;/span&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;12&lt;/span&gt;&lt;span&gt;ids, docs, metas &lt;span style="color:#f92672"&gt;=&lt;/span&gt; [], [], []
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;13&lt;/span&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; it &lt;span style="color:#f92672"&gt;in&lt;/span&gt; items:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;14&lt;/span&gt;&lt;span&gt; guid &lt;span style="color:#f92672"&gt;=&lt;/span&gt; (it&lt;span style="color:#f92672"&gt;.&lt;/span&gt;findtext(&lt;span style="color:#e6db74"&gt;&amp;#34;guid&amp;#34;&lt;/span&gt;) &lt;span style="color:#f92672"&gt;or&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;strip()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;15&lt;/span&gt;&lt;span&gt; title &lt;span style="color:#f92672"&gt;=&lt;/span&gt; (it&lt;span style="color:#f92672"&gt;.&lt;/span&gt;findtext(&lt;span style="color:#e6db74"&gt;&amp;#34;title&amp;#34;&lt;/span&gt;) &lt;span style="color:#f92672"&gt;or&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;strip()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;16&lt;/span&gt;&lt;span&gt; desc &lt;span style="color:#f92672"&gt;=&lt;/span&gt; re&lt;span style="color:#f92672"&gt;.&lt;/span&gt;sub(&lt;span style="color:#e6db74"&gt;r&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;\s+&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34; &amp;#34;&lt;/span&gt;, re&lt;span style="color:#f92672"&gt;.&lt;/span&gt;sub(&lt;span style="color:#e6db74"&gt;r&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&amp;lt;[^&amp;gt;]+&amp;gt;&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34; &amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;17&lt;/span&gt;&lt;span&gt; it&lt;span style="color:#f92672"&gt;.&lt;/span&gt;findtext(&lt;span style="color:#e6db74"&gt;&amp;#34;description&amp;#34;&lt;/span&gt;) &lt;span style="color:#f92672"&gt;or&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;))&lt;span style="color:#f92672"&gt;.&lt;/span&gt;strip()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;18&lt;/span&gt;&lt;span&gt; ids&lt;span style="color:#f92672"&gt;.&lt;/span&gt;append(guid)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;19&lt;/span&gt;&lt;span&gt; docs&lt;span style="color:#f92672"&gt;.&lt;/span&gt;append(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;title&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\n\n&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;desc&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#f92672"&gt;.&lt;/span&gt;strip())
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;20&lt;/span&gt;&lt;span&gt; metas&lt;span style="color:#f92672"&gt;.&lt;/span&gt;append({&lt;span style="color:#e6db74"&gt;&amp;#34;link&amp;#34;&lt;/span&gt;: (it&lt;span style="color:#f92672"&gt;.&lt;/span&gt;findtext(&lt;span style="color:#e6db74"&gt;&amp;#34;link&amp;#34;&lt;/span&gt;) &lt;span style="color:#f92672"&gt;or&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;strip()})
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;21&lt;/span&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;22&lt;/span&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; ids:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;23&lt;/span&gt;&lt;span&gt; col&lt;span style="color:#f92672"&gt;.&lt;/span&gt;upsert(ids&lt;span style="color:#f92672"&gt;=&lt;/span&gt;ids, documents&lt;span style="color:#f92672"&gt;=&lt;/span&gt;docs, metadatas&lt;span style="color:#f92672"&gt;=&lt;/span&gt;metas)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;24&lt;/span&gt;&lt;span&gt;col&lt;span style="color:#f92672"&gt;.&lt;/span&gt;modify(metadata&lt;span style="color:#f92672"&gt;=&lt;/span&gt;{&lt;span style="color:#e6db74"&gt;&amp;#34;last_sync&amp;#34;&lt;/span&gt;: int(time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time())})
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;One caveat: ChromaDB does not really work when you query big data - feed entries
fit in one record, but a whole crawled page would end up as one giant embedding
that drowns the details. So chunk bigger records: split the text into blocks of
1024 or 4096 characters, upsert one record per block, id like
&lt;code&gt;f&amp;quot;{url}#{offset}&amp;quot;&lt;/code&gt;, URL in the metadata - and every hit points straight back
to its source.&lt;/p&gt;
&lt;p&gt;Searching it later is a one-liner:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;1&lt;/span&gt;&lt;span&gt;r &lt;span style="color:#f92672"&gt;=&lt;/span&gt; col&lt;span style="color:#f92672"&gt;.&lt;/span&gt;query(query_texts&lt;span style="color:#f92672"&gt;=&lt;/span&gt;[&lt;span style="color:#e6db74"&gt;&amp;#34;hugo template lookup order&amp;#34;&lt;/span&gt;], n_results&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;10&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;2&lt;/span&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; iid, m, d &lt;span style="color:#f92672"&gt;in&lt;/span&gt; zip(r[&lt;span style="color:#e6db74"&gt;&amp;#34;ids&amp;#34;&lt;/span&gt;][&lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;], r[&lt;span style="color:#e6db74"&gt;&amp;#34;metadatas&amp;#34;&lt;/span&gt;][&lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;], r[&lt;span style="color:#e6db74"&gt;&amp;#34;distances&amp;#34;&lt;/span&gt;][&lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;]):
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;3&lt;/span&gt;&lt;span&gt; print(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;d&lt;span style="color:#e6db74"&gt;:&lt;/span&gt;&lt;span style="color:#e6db74"&gt;.3f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt; | &lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;m[&lt;span style="color:#e6db74"&gt;&amp;#39;link&amp;#39;&lt;/span&gt;]&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Add structured filters on the metadata you stored - &lt;code&gt;where={&amp;quot;published&amp;quot;: {&amp;quot;$gte&amp;quot;: since}}&lt;/code&gt;
to narrow by date, or &lt;code&gt;where={&amp;quot;tags&amp;quot;: &amp;quot;hugo&amp;quot;}&lt;/code&gt; if you tagged each page - instead of
trusting the vectors alone.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;OpenSearch&lt;/strong&gt; (or Elasticsearch, or a similar Lucene-based engine) - when you
outgrow the simple case. Real BM25 scoring, filters, aggregations, dashboards. It
runs fine as a single-node container:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;1&lt;/span&gt;&lt;span&gt;podman container run --name opensearch-db &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;2&lt;/span&gt;&lt;span&gt; -e &lt;span style="color:#e6db74"&gt;&amp;#34;discovery.type=single-node&amp;#34;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;3&lt;/span&gt;&lt;span&gt; -e OPENSEARCH_INITIAL_ADMIN_PASSWORD&lt;span style="color:#f92672"&gt;=&lt;/span&gt;SECRET_HERE &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;4&lt;/span&gt;&lt;span&gt; -p 9200:9200 -p 9300:9300 &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;5&lt;/span&gt;&lt;span&gt; -v opensearch_data:/usr/share/opensearch &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;6&lt;/span&gt;&lt;span&gt; docker.io/opensearchproject/opensearch
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Index the harvest as records too - page content in blocks of 1024 or 4096
characters, the source URL and any tags as plain fields - and search full text
over &lt;code&gt;content&lt;/code&gt;, no vectors involved. More ceremony than the ChromaDB one-liner:
the official client, plus the dev container&amp;rsquo;s self-signed TLS:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 1&lt;/span&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; opensearchpy &lt;span style="color:#f92672"&gt;import&lt;/span&gt; OpenSearch
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 2&lt;/span&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 3&lt;/span&gt;&lt;span&gt;client &lt;span style="color:#f92672"&gt;=&lt;/span&gt; OpenSearch(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 4&lt;/span&gt;&lt;span&gt; hosts&lt;span style="color:#f92672"&gt;=&lt;/span&gt;[{&lt;span style="color:#e6db74"&gt;&amp;#34;host&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;127.0.0.1&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;port&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;9200&lt;/span&gt;}],
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 5&lt;/span&gt;&lt;span&gt; http_auth&lt;span style="color:#f92672"&gt;=&lt;/span&gt;(&lt;span style="color:#e6db74"&gt;&amp;#34;admin&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;SECRET_HERE&amp;#34;&lt;/span&gt;),
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 6&lt;/span&gt;&lt;span&gt; use_ssl&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#66d9ef"&gt;True&lt;/span&gt;, verify_certs&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#66d9ef"&gt;False&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 7&lt;/span&gt;&lt;span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 8&lt;/span&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt; 9&lt;/span&gt;&lt;span&gt;r &lt;span style="color:#f92672"&gt;=&lt;/span&gt; client&lt;span style="color:#f92672"&gt;.&lt;/span&gt;search(index&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;pages&amp;#34;&lt;/span&gt;, body&lt;span style="color:#f92672"&gt;=&lt;/span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;10&lt;/span&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;size&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;10&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;11&lt;/span&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;query&amp;#34;&lt;/span&gt;: {&lt;span style="color:#e6db74"&gt;&amp;#34;bool&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;12&lt;/span&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;must&amp;#34;&lt;/span&gt;: {&lt;span style="color:#e6db74"&gt;&amp;#34;match&amp;#34;&lt;/span&gt;: {&lt;span style="color:#e6db74"&gt;&amp;#34;content&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;hugo template lookup order&amp;#34;&lt;/span&gt;}},
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;13&lt;/span&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;filter&amp;#34;&lt;/span&gt;: {&lt;span style="color:#e6db74"&gt;&amp;#34;term&amp;#34;&lt;/span&gt;: {&lt;span style="color:#e6db74"&gt;&amp;#34;tags&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;hugo&amp;#34;&lt;/span&gt;}},
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;14&lt;/span&gt;&lt;span&gt; }},
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;15&lt;/span&gt;&lt;span&gt;})
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;16&lt;/span&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; hit &lt;span style="color:#f92672"&gt;in&lt;/span&gt; r[&lt;span style="color:#e6db74"&gt;&amp;#34;hits&amp;#34;&lt;/span&gt;][&lt;span style="color:#e6db74"&gt;&amp;#34;hits&amp;#34;&lt;/span&gt;]:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span style="white-space:pre;-webkit-user-select:none;user-select:none;margin-right:0.4em;padding:0 0.4em 0 0.4em;color:#7f7f7f"&gt;17&lt;/span&gt;&lt;span&gt; print(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#39;&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;hit[&lt;span style="color:#e6db74"&gt;&amp;#34;_score&amp;#34;&lt;/span&gt;]&lt;span style="color:#e6db74"&gt;:&lt;/span&gt;&lt;span style="color:#e6db74"&gt;.3f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt; | &lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;hit[&lt;span style="color:#e6db74"&gt;&amp;#34;_source&amp;#34;&lt;/span&gt;][&lt;span style="color:#e6db74"&gt;&amp;#34;link&amp;#34;&lt;/span&gt;]&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#39;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Same shape as the ChromaDB loop, with the BM25 &lt;code&gt;_score&lt;/code&gt; in place of similarity
distance. Drop the &lt;code&gt;filter&lt;/code&gt; clause for plain full text; the &lt;code&gt;term&lt;/code&gt; query on tags
only behaves when the field is mapped as &lt;code&gt;keyword&lt;/code&gt; - not analyzed into separate
words. And when &lt;code&gt;bool&lt;/code&gt; is not enough, the whole query DSL lives in the &lt;code&gt;body&lt;/code&gt;:
&lt;code&gt;match_phrase&lt;/code&gt; for word order, &lt;code&gt;range&lt;/code&gt; for dates, aggregations for counts.&lt;/p&gt;
&lt;p&gt;For &amp;ldquo;I want to find that blog post I half-remember&amp;rdquo;, ChromaDB is honestly enough.
Reach for OpenSearch when you want query languages, exact-term matching at scale,
or you are indexing millions of documents. There is also the middle ground I use
daily: &lt;code&gt;ripgrep&lt;/code&gt; over plain markdown files when the corpus is just files on disk.&lt;/p&gt;
&lt;h1 id="what-is-still-missing---the-honest-part"&gt;What is still missing - the honest part&lt;/h1&gt;
&lt;p&gt;Let me not oversell this. What the DIY pipeline does not give you:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;No global relevance ranking.&lt;/strong&gt; Google&amp;rsquo;s index spans the whole web and ranks it.
My pipeline answers &amp;ldquo;what does &lt;em&gt;this&lt;/em&gt; site have&amp;rdquo; - you need to know the domain
(or a seed page) first. Discovery across the whole web is a different, harder game.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stale or missing data.&lt;/strong&gt; Sitemaps lie (or rather, lag). Not every site has
feeds. Some sitemaps are generated once a year and never touched again.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;JavaScript-only sites.&lt;/strong&gt; SPA shells return an empty &lt;code&gt;&amp;lt;div id=&amp;quot;root&amp;quot;&amp;gt;&lt;/code&gt; to curl.
Headless chromium in dump-DOM mode (&lt;code&gt;--dump-dom --virtual-time-budget=10000&lt;/code&gt;)
renders them, at the cost of a heavy page load - use sparingly, and throttle.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;You are the ranking function.&lt;/strong&gt; The agent reading ten titles and summaries and
picking the right one is, effectively, doing the ranking. That works remarkably
well in practice - but it is a different trade-off than a tuned index.&lt;/li&gt;
&lt;/ul&gt;
&lt;h1 id="what-is-left"&gt;What is left?&lt;/h1&gt;
&lt;p&gt;The pipeline above answers &amp;ldquo;where on this site is X&amp;rdquo;. A few problems it deliberately
leaves open - and where I think the interesting work is:&lt;/p&gt;
&lt;h2 id="rankings"&gt;Rankings&lt;/h2&gt;
&lt;p&gt;What we do not have is PageRank. Google&amp;rsquo;s relevance is built on the link graph - who
points at whom, weighted and iterated - plus decades of click-through and quality
signals. Locally we have none of that. But do we actually need it? A per-site corpus
is small; a hundred titles fit in an agent&amp;rsquo;s context window without any ranking at
all. Where it starts to matter is choosing &lt;em&gt;between&lt;/em&gt; sources, and there are cheap
proxies that go a long way:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;recency - feeds give it for free,&lt;/li&gt;
&lt;li&gt;your own hit rate - did this source actually answer anything last time? Keep the
score per domain and you get a &lt;em&gt;personal&lt;/em&gt; relevance signal no crawler can fake,&lt;/li&gt;
&lt;li&gt;curated seeds - the sites you bookmarked once and keep coming back to.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And when none of that settles it, the agent reads ten titles and picks. That works
remarkably well - but every such judgment is an AI call, so caching the verdicts
(&amp;ldquo;this blog is a good source for X&amp;rdquo;) is where the real savings live.&lt;/p&gt;
&lt;h2 id="crawl-only-what-is-worth-crawling"&gt;Crawl only what is worth crawling&lt;/h2&gt;
&lt;p&gt;We do not need all the world&amp;rsquo;s slop. A crawl budget - even a generous one - should
follow worthiness, not completeness. The signals are the same ones as above, applied
&lt;em&gt;before&lt;/em&gt; spending any request:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;does the site &lt;em&gt;bother&lt;/em&gt; to publish a sitemap and feeds? That alone is a statement
of intent: it wants to be read by machines. Sites without any of it force you into
page-by-page guessing - the most expensive kind of crawl - and often are not worth
the guess,&lt;/li&gt;
&lt;li&gt;past usefulness - a domain whose pages answered your questions five times earns
more budget than one that answered zero,&lt;/li&gt;
&lt;li&gt;who references it - your own notes, bookmarks and previous answers are a link
graph, just a private one.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Update the per-domain score every time a fetch proves useful or useless and you end
up with a tiny, self-correcting priority list. Effectively a PageRank where the
&amp;ldquo;links&amp;rdquo; are your own queries that led somewhere.&lt;/p&gt;
&lt;h2 id="recrawl-policies"&gt;Recrawl policies&lt;/h2&gt;
&lt;p&gt;The last open piece: when to come back. Feeds are almost self-solving - poll them,
use conditional requests (&lt;code&gt;ETag&lt;/code&gt;, &lt;code&gt;If-Modified-Since&lt;/code&gt;), and an unchanged site costs
you a 304. Sitemaps are trickier: how often to re-fetch depends on the observed
change rate, which you get for free from the &lt;code&gt;&amp;lt;lastmod&amp;gt;&lt;/code&gt; diffs mentioned above - a
news site and a personal blog clearly deserve different schedules. There is also the
question of expiring stale entries from the local index versus keeping them as
&amp;ldquo;was true once&amp;rdquo; history. Each of these deserves more space than a bullet - maybe
the subject of the next post.&lt;/p&gt;
&lt;h1 id="so---do-we-need-google"&gt;So - do we need Google?&lt;/h1&gt;
&lt;p&gt;For finding &lt;em&gt;anything on the whole web&lt;/em&gt; - probably yes, for now. The long tail of
discovery is what the big indexes are genuinely good at.&lt;/p&gt;
&lt;p&gt;But for the question I actually ask dozens of times a day - &lt;em&gt;&amp;ldquo;where on this site is
the thing I remember?&amp;rdquo;&lt;/em&gt; - the answer is: we do not need a search engine billing us
per query. &lt;code&gt;robots.txt&lt;/code&gt;, a sitemap and a feed are a site&amp;rsquo;s own index, published for
free, machine-readable, waiting to be used. A skill that reads them in order and a
local ChromaDB to keep the harvest replace a surprisingly large chunk of what we
have been paying for. And unlike a commercial API, the free meal comes with an
etiquette: read &lt;code&gt;robots.txt&lt;/code&gt;, honor &lt;code&gt;Crawl-delay&lt;/code&gt;, prefer sitemaps and feeds over
page-by-page crawling.&lt;/p&gt;
&lt;p&gt;One last thing, since we are all someone else&amp;rsquo;s crawler target: publish your own
sitemap and feeds. This very site does - Hugo generates &lt;code&gt;/sitemap.xml&lt;/code&gt; and
&lt;code&gt;/index.xml&lt;/code&gt; (RSS) and &lt;code&gt;/atom.xml&lt;/code&gt; out of the box, zero extra work. Do it, and the
agents of the world can find your content without paying a middleman for the
privilege.&lt;/p&gt;</content></entry></feed>