Crawler policy

About DRI-bot

DRI-bot is the crawler operated by DRI, the persistent-identifier service for research publications. It reads journal and repository websites so that new articles can be assigned a permanent identifier, and so that identifiers already issued keep pointing at a page that is still there.

How to recognise it

Every request carries this User-Agent:

DRI-bot/1.0 (+https://dri-id.everedgegroup.com/bot; mailto:ops@dri-id.everedgegroup.com)

Requests are made over HTTPS only — DRI-bot never fetches an http:// URL, and never downgrades to one on a redirect.

What it fetches, and how often

Responses are size- and time-capped, and requests are not parallelised across a single host beyond what a feed's own contents require.

robots.txt

DRI-bot honours robots.txt, reading the rules for the DRI-bot token before it reads anything else on the host. To keep it off your site entirely:

User-agent: DRI-bot
Disallow: /

Or to keep it out of one area while leaving articles readable:

User-agent: DRI-bot
Disallow: /admin/
Disallow: /members-only/

A disallowed page is simply not fetched. Identifiers already issued for it keep resolving — blocking the crawler withdraws nothing.

Allowing it through bot protection

Cloudflare Bot Fight Mode, Wordfence, Sucuri and similar tools challenge unknown crawlers, which DRI-bot cannot answer: it does not execute JavaScript and does not solve challenges. A challenged request looks to us like a site that has stopped publishing.

Allow it by the User-Agent string above — in Cloudflare, a WAF custom rule such as http.user_agent contains "DRI-bot" with the action set to Skip; in Wordfence, add DRI-bot to the allowlisted crawlers.

Contact

Questions, complaints, or a request to stop crawling: ops@dri-id.everedgegroup.com. We answer from a person, and "please stop" needs no justification.