Identification
User-Agent: worldir-crawler/0.1. Requests include crawler@worldir.net as the operator contact.
Responsible crawling
Our crawler builds a source-linked engineering index for human and agent workflows. It identifies itself as worldir-crawler and follows explicit source boundaries.
User-Agent: worldir-crawler/0.1. Requests include crawler@worldir.net as the operator contact.
Public jobs honor robots.txt, apply per-domain pacing and backoff, cap response sizes, and refuse local/private network targets.
Every crawl is restricted to reviewed hosts and URL subtrees. Cross-host redirects require a separately governed source seed.
Results retain the source URL, observation snapshot, immutable asset id and reviewed terms URL. Unknown public rights default to index/search only.