<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AI Crawlers on Jafforge Blog</title><link>https://jafforge.com/tags/ai-crawlers/</link><description>Recent content in AI Crawlers on Jafforge Blog</description><generator>Hugo -- 0.146.0</generator><language>en-us</language><lastBuildDate>Mon, 28 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://jafforge.com/tags/ai-crawlers/index.xml" rel="self" type="application/rss+xml"/><item><title>Building webtelemetry.dev: A Live Survey of the Web as It Launches</title><link>https://jafforge.com/posts/webtelemetry-live-survey-of-new-websites/</link><pubDate>Mon, 28 Sep 2026 00:00:00 +0000</pubDate><guid>https://jafforge.com/posts/webtelemetry-live-survey-of-new-websites/</guid><description>Why and how I built webtelemetry.dev: a live survey of new websites, and my way of learning how to process large amounts of data with Kafka and PostgreSQL. Plus what it found about AI crawlers and llms.txt.</description><content:encoded><![CDATA[<p>New websites go live every few seconds. Someone registers a domain, points it at a
server, gets a certificate and publishes a first page. I wanted to see that
happening: what people use to build new sites today, where they host them, and
whether they let AI crawlers read them.</p>
<p>To be honest, the websites were only part of the reason. I mostly built this to
learn how to process large amounts of data. I wanted to work with a stream of data
that never stops, get real experience with tools like Kafka, and see what it takes
to keep a pipeline running when millions of records come in every day. I also
wanted to build a big database that I can use for my next projects. New websites
turned out to be a great subject for that: the data is public, there is always
more of it, and there are interesting questions to answer.</p>
<p>So I built <strong>webtelemetry.dev</strong>, with a lot of help from Claude. It is a live survey
of new websites: it finds new sites in public certificate logs, checks with each
domain&rsquo;s registry that the domain is really new, visits the homepage once, and
publishes what it finds: live statistics, a profile for every site, and a free
dataset every day.</p>
<p><a href="https://webtelemetry.dev/">Visit webtelemetry.dev</a> · <a href="/projects/webtelemetry/">Explore the project</a> · <a href="https://webtelemetry.dev/data">Download the data</a></p>
<h2 id="what-it-shows">What it shows</h2>
<p>My favourite part is the live map. Every dot is a newly registered domain whose
server location we know. Dots appear as soon as the crawler confirms the domain,
and the newest sites are listed next to the map.</p>
<p><img alt="Live map of new websites appearing around the world, with the newest sites listed on the right" loading="lazy" src="/posts/webtelemetry-live-survey-of-new-websites/live-map.jpg"></p>
<p>Every site gets its own profile: what it is built with, where it is hosted, who
runs its DNS and email, when the domain was registered, how fast it responded and
which AI crawlers it blocks. New sites that pass a basic quality check also get a
screenshot of their homepage.</p>
<p><img alt="Site profile for a new SaaS product, with its homepage screenshot, detected stack and hosting details" loading="lazy" src="/posts/webtelemetry-live-survey-of-new-websites/site-profile.jpg"></p>
<p>The statistics pages add all of this up: which platforms and frameworks new sites
use, which countries and networks host them, who provides their DNS and email, and
which certificate authorities and registrars they use. Every technology, country
and type of site also has its own page.</p>
<p><img alt="What websites are built with: PHP, WordPress, Node.js, React, Next.js and more, with a trend line for each" loading="lazy" src="/posts/webtelemetry-live-survey-of-new-websites/built-with.png"></p>
<h2 id="finding-new-websites-when-there-is-no-list-of-new-websites">Finding new websites when there is no list of new websites</h2>
<p>There is no public list of &ldquo;websites launched today&rdquo;. But there is something close:
<strong>Certificate Transparency</strong>. Every TLS certificate issued by a public authority is
written to public logs that browsers require, and a new website almost always gets a
certificate within minutes of going live.</p>
<p>There are two problems with that. First, most certificates are renewals for sites
that have existed for years. Second, the logs are huge: the busiest ones get well
over a hundred new entries every second. When I read them in order, the crawler was
almost a day behind after its first week. So now it reads only the newest batch from
each of 17 logs, run by five different companies, every 10 seconds. That still adds
up to around 20 million certificates a day. It is a live sample, not a full copy,
and when it falls behind it skips ahead to the newest entries instead of catching
up. For this project, being late is worse than missing a few sites.</p>
<p>To separate new sites from renewals, every domain goes through a check first. The
crawler asks the domain&rsquo;s own registry, over RDAP (the modern replacement for
WHOIS), when the domain was registered:</p>
<ul>
<li>domains registered in the last 30 days are visited first,</li>
<li>domains up to a year old come next,</li>
<li>and a random 2% of older domains are visited as a comparison sample.</li>
</ul>
<p>Registries are shared infrastructure, so the crawler spaces out its requests to each
one. Verisign, which runs .com and .net, gets at most one request per second, in
line with its terms of use.</p>
<h2 id="how-the-pipeline-works">How the pipeline works</h2>
<p><img alt="How webtelemetry.dev works: certificate logs feed a registration check and a queue; a fetcher, parser and sink turn pages into rows in PostgreSQL; the website, open dataset and screenshots are built from there" loading="lazy" src="/posts/webtelemetry-live-survey-of-new-websites/architecture.png"></p>
<p>Everything runs on one server with 8 cores, using Docker Compose. The parts talk to
each other through Kafka topics:</p>
<ul>
<li><strong>Discovery</strong> reads the certificate logs and sends out candidate domains.</li>
<li><strong>The gate</strong> looks up registration dates and decides which domains are worth
visiting.</li>
<li><strong>The scheduler</strong> hands out visits from a queue in Postgres. New domains always go
first, before any revisits.</li>
<li><strong>The fetcher</strong> reads robots.txt first and respects it, then fetches the homepage,
up to three of the site&rsquo;s own scripts and <code>/llms.txt</code>. It waits at least two
seconds between requests to the same host.</li>
<li><strong>The parser</strong> turns the page into a profile.</li>
<li><strong>The sink</strong> saves one row per site, plus a history of what changed.</li>
</ul>
<p>Using Kafka on a single server might look like overkill, but learning it was one of
the reasons I started this project, and it has earned its place. When one part
slows down, messages simply wait in Kafka instead of getting lost, and I can restart
or scale any part without losing work. It also taught me a lesson the hard way:
KafkaJS processes one message per partition at a time, so increasing the fetcher&rsquo;s
concurrency did nothing until the topic had enough partitions.</p>
<h2 id="turning-a-homepage-into-a-profile">Turning a homepage into a profile</h2>
<p>The parser checks each page against about 7,500 technology fingerprints from two
open rule sets, both GPL-3.0 forks of Wappalyzer:
<a href="https://github.com/HTTPArchive/wappalyzer">HTTP Archive&rsquo;s</a> and
<a href="https://github.com/enthec/webappanalyzer">enthec&rsquo;s webappanalyzer</a>. For about a
hundred of the most important technologies I wrote my own rules, tuned against
false positives I found. Around 6,600 technologies can be detected from a single
visit to the homepage; the rest need a real browser.</p>
<p>Adding the second rule set taught me not to trust data blindly. Before turning it
on, I ran both rule sets over 1,869 real homepages and went through every new match
by hand. Ten rules were matching the wrong things. One would have marked every
Shopify store as using a JavaScript physics engine, because its pattern also matched
a Shopify theme file called <code>rte-formatter.js</code>. Another group looked for the
verification records that services like Slack or OpenAI ask domain owners to add,
and turned them into claims like &ldquo;this site uses OpenAI&rdquo;. I left those rules out,
and I run the same check before every update of the rule sets.</p>
<p>Besides technologies, each profile records the hosting network and country, DNS and
email providers, SPF and DMARC, IPv6, the certificate issuer, a quality score and a
guess at what type of site it is. Adult, drug, weapons, piracy and scam sites are
removed completely. For that the crawler uses the page&rsquo;s own text, Cloudflare&rsquo;s
family DNS filter and an image classifier that checks the screenshot.</p>
<h2 id="ai-crawlers-and-llmstxt">AI crawlers and llms.txt</h2>
<p>The crawler has to read robots.txt anyway, to know if it is allowed to visit. The
same file also says which AI crawlers a site blocks, so every visit records the rules
for fifteen of them, including GPTBot, ClaudeBot, CCBot, Google-Extended and
PerplexityBot.</p>
<p><img alt="AI crawlers page: the 10,000 most popular websites block AI crawlers far more often than new websites" loading="lazy" src="/posts/webtelemetry-live-survey-of-new-websites/ai-crawlers.png"></p>
<p>As of 28 September 2026, 12.9% of the 10,000 most popular websites block at least
one AI crawler from the whole site, compared with 2.1% of new websites. So popular
sites block AI crawlers about six times as often. The data is still young and the
numbers change every day, but this gap has been there since the first hours.</p>
<p>My first version of this comparison was wrong, and I think the mistake is worth
sharing. I compared new sites with &ldquo;all other sites&rdquo;, and it turned out that most of
those other sites came from the Tranco top million, which is a list of popular
sites. When I compared new sites with a random sample of ordinary older sites
instead, the big difference disappeared: the two are within about one percentage
point of each other. The real difference was popularity, not age. Now every
comparison uses clearly named groups (new sites, a random sample of older sites, the
top 10,000 and the rest of the top million), and the site only reports a difference
when a statistical test says it is unlikely to be chance. Blocking is also split in
two: crawlers blocked from the whole site, and crawlers blocked from only some
pages.</p>
<p><code>llms.txt</code> is the opposite signal: a plain text file that tells language models what
a site is about and which pages matter. About one in eight new websites has one, but
many of those files were not written by the site owners at all.</p>
<p><img alt="Who wrote the llms.txt files of new websites: the platform, an SEO plugin, a shared template, or no known tool" loading="lazy" src="/posts/webtelemetry-live-survey-of-new-websites/llms-txt-makers.png"></p>
<p>Shopify and Wix create one for every site on their platforms, SEO plugins generate
them from a site&rsquo;s pages, and parking pages all share the same template. The crawler
recognises the lines these tools leave in their files: almost a fifth of the
llms.txt files on new sites came from a platform or a plugin, and about a quarter
are templates shared with other sites.</p>
<h2 id="fast-open-and-easy-to-cite">Fast, open and easy to cite</h2>
<p>Every page reads from summary tables (materialized views that Postgres refreshes in
the background every few minutes), so pages stay fast no matter how big the data
gets. The frontend is built with Astro and rendered on the server, and the live
counters and the map update in real time using server-sent events. There are no
accounts and no advertising trackers, and the fonts are self-hosted.</p>
<p>All the data is meant to be reused:</p>
<ul>
<li>a daily CSV of every site, free under CC BY 4.0, with a snapshot kept every week;</li>
<li>CSV downloads for the main charts, and a JSON API;</li>
<li>a &ldquo;Cite this&rdquo; box on every chart, with the date of the figures.</li>
</ul>
<p><img alt="Cite this box open on a chart: a ready citation with the date and a permanent link" loading="lazy" src="/posts/webtelemetry-live-survey-of-new-websites/cite-this.png"></p>
<p>Two more features will fill in as the data gets older. Every Monday, starting 5
October, 1,000 launches found one or two weeks earlier are checked against their
registry again, and the results are published as a CSV. And the survival figures
will show how many of each week&rsquo;s launches are still working after one, four, eight
and twelve weeks.</p>
<h2 id="running-it-on-one-server">Running it on one server</h2>
<p>The whole stack is Kafka 4, PostgreSQL 17, Node 22 with TypeScript, Astro 7 and a
Playwright container for screenshots, behind Caddy and Cloudflare, on one server
with 8 cores and 24 GB of RAM. A deploy script builds each service, runs a quick test
and rolls back automatically if the new version does not start crawling. There are
nightly backups, alerts and a small status bot on Telegram, and the admin panel is
protected with a password and two-factor authentication.</p>
<h2 id="built-with-claude">Built with Claude</h2>
<p>I did not build this alone. Claude helped me build most of it: I worked with it in
Claude Code, where I described what I wanted and it wrote most of the code with me,
from the certificate log reader and the Kafka pipeline to the website and the admin
panel. It also did the full optimisation work, which is where I learned the most:</p>
<ul>
<li>moving every page onto summary tables once the database got big, so pages stay
fast;</li>
<li>tuning Kafka partitions and batching, and Postgres for the bigger server;</li>
<li>the registration check, with a separate queue and rate limit for every registry;</li>
<li>testing the second rule set on real homepages before switching it on;</li>
<li>a deploy script that runs a smoke test and rolls back on its own.</li>
</ul>
<p>My part was deciding what to build, testing it, and asking for changes whenever
something did not look right.</p>
<h2 id="what-i-learned">What I learned</h2>
<ol>
<li><strong>Working with a lot of data is mostly about the boring parts.</strong> Summary tables,
Kafka partitions, disk space and backups mattered more than any clever query. The
hard part is keeping everything fast and reliable while the data keeps growing.</li>
<li><strong>For this project, fresh data beats complete data.</strong> Reading only the newest
certificates tells you more about what is launching right now than reading
everything and being a day behind.</li>
<li><strong>Check what you are comparing, not just the numbers.</strong> My first result about AI
crawlers compared the wrong groups of sites. Now every group has a clear name, and
every difference is tested before the site reports it.</li>
<li><strong>Test other people&rsquo;s rules on your own data.</strong> Open rule sets are a gift, but ten
of these rules would have quietly messed up my numbers.</li>
<li><strong>Be polite to shared infrastructure.</strong> Registries, certificate logs and the
websites themselves are shared resources. Rate limits, respecting robots.txt and a
clear page about the bot cost very little and keep the project welcome.</li>
</ol>
<p>webtelemetry.dev is live and collecting data every minute. The numbers will get more
reliable every week, and in early October the first seven-day average of launches
and the first launch audit will appear. It is also turning into the big database I
wanted for my next projects: every site it checks adds to a growing record of how
the web is built. If you work with web data, the daily dataset is free to use.</p>
<p><a href="https://webtelemetry.dev/">Visit webtelemetry.dev</a> · <a href="/projects/webtelemetry/">Explore the project</a> · <a href="https://webtelemetry.dev/data">Download the data</a></p>
]]></content:encoded></item></channel></rss>