Skip to content

MDUBot

MDUBot indexes .mdu.pm sites for MDU Search. It identifies itself as:

MDUBot/1.0 (+https://name.mdu.pm/bot)

What it visits

Only sites whose owner opted in. Registering a name does not put you in the index - a holder has to switch on their directory listing, and switching it off removes their pages immediately.

By default it reads one page: the front page. A holder may name specific paths, and then the pages linked directly from those are read too. If the site publishes a sitemap, the URLs it lists are read as well - that is how a site with hundreds of articles gets indexed properly rather than by guesswork.

It stays on the site, with one exception: where a .mdu.pm name permanently redirects to content the registry also operates, the destination is read and credited to both hosts. The destination's own robots.txt still decides.

Keeping it out

robots.txt is fully respected, including a group naming MDUBot in preference to the wildcard group. So are <meta name="robots" content="noindex"> and rel="nofollow". To exclude it entirely:

User-agent: MDUBot
Disallow: /

How it behaves

It waits at least a second between requests, honors Crawl-delay, and reads at most 40 pages per site per run. A site that has not changed is not re-read.

Something wrong? Report it.