Replying to
@felwert@fedihum.org Not sure if it's best practice or just old-school: honest user agent (name of the software, version number, possibly the harvesting purpose or person responsible), check robots.txt for delay/interval instructions before harvesting, use generous delays/intervals between requests if no instructions are given, and never fire async requests.
This is how parts of the data for the NFDI4Culture Knowledge Graph are harvested 🤖