WORLDIROpen console

Responsible crawling

WorldIR collects governed public engineering evidence.

Our crawler builds a source-linked engineering index for human and agent workflows. It identifies itself as worldir-crawler and follows explicit source boundaries.

Identification

User-Agent: worldir-crawler/0.1. Requests include crawler@worldir.net as the operator contact.

Access policy

Public jobs honor robots.txt, apply per-domain pacing and backoff, cap response sizes, and refuse local/private network targets.

Scope

Every crawl is restricted to reviewed hosts and URL subtrees. Cross-host redirects require a separately governed source seed.

Evidence and rights

Results retain the source URL, observation snapshot, immutable asset id and reviewed terms URL. Unknown public rights default to index/search only.