About our crawler
When we research a school's housing options we read the school's own public pages — the residence descriptions, the fees, the application rules — so that what we tell students is what the school says. The requests come from a crawler that identifies itself as DecidientBot with a link to this page.
How it behaves
- It reads robots.txt on every site first and keeps to it. A page your robots.txt asks crawlers not to read is not read, and we do not look for a copy of it elsewhere.
- It makes one request at a time to a site, a few seconds apart, and reads a few dozen pages per school at most. It never crawls a site whole.
- If a page will not load — a bot check, a rate limit, a timeout — it stops and tries once more hours later. It does not try to get past a bot check, and it uses no proxies or disguised browsers.
- What it reads is used to write cited facts for the decision-support site about that school. Every fact links back to the page it came from.
To opt out or ask a question
Add a Disallow rule for DecidientBot (or any of the names it used before: L3Bot, LearnLoopBot or LearnLoopLiveBot) to your robots.txt, or write to [email protected] and we will stop reading the pages you name. We answer within a few days.