About our crawler
Operated by MA EdTech Solutions Inc. · crawler@educationall.tech · last updated 2026-09-27
If you found AIWikiBot in your server logs, this page is for you.
The short version
We fetch publicly available pages from websites of organizations that serve children and families, to build a directory of those services. We identify ourselves honestly, we obey robots.txt, we fetch slowly, and we stop when asked. If you want us to stop, email crawler@educationall.tech and we will — permanently. You do not need to explain why.
Who we are
AIWikiBot is operated by MA EdTech Solutions Inc. Our contact address is crawler@educationall.tech, and it is read by a person.
Our User-Agent is:
AIWikiBot/0.1 (+https://bot.educationall.tech; crawler@educationall.tech)
We do not disguise this. We never present ourselves as a browser, and we do not use any technique to make our requests harder to identify. If you block us, we stay blocked.
What we are building
A directory of organizations that serve children and families in Ontario — paediatric occupational therapy, speech-language pathology, clinics, community centres and similar services — so that families and the professionals who support them can find them by what they need rather than by guessing a search term.
We began in Waterloo Region and widened to the province on 2026-08-07. If you are reading this because our crawler reached a site outside Waterloo Region, that widening is why, and this sentence was updated in the same change.
We record factual, publicly published details: an organization's name, address, phone number, website, and the services it says it offers. Every one of those facts is stored with a link to the exact page we read it from and the date we read it, so anything we publish can be traced back to your own words.
What we do not do
These are constraints in our software, not intentions.
- We never log in. We do not create accounts, and we never submit a form.
- We never try to get past a paywall, a login wall, or a CAPTCHA. If we encounter one, we record that we were blocked and move on.
- We do not collect personal information about your clients or patients. If a page appears to contain client information — an intake form, a testimonial naming a child, a directory listing that was not meant to be public — our software discards the page without storing it, and records only that this happened.
- We do not collect personal email addresses, and we do not use anything we fetch to build a mailing list.
- We do not copy your content. We record facts and short quotations as evidence for those facts. We do not reproduce your pages, and we do not republish your text.
How often we visit, and how gently
- One request at a time. We never open parallel connections to your site.
- We wait between requests — at least one second, and normally about ten times however long your server took to answer, up to thirty seconds. A slower server gets a slower crawler.
- We honour
Retry-After, and we back off automatically on 429 and 503.
- We stop after three consecutive refusals, for at least a week, without a human needing to notice.
- We fetch a limited number of pages, which depends on the kind of organization. From a small organization, a funder or a business, at most ten. From a network's own site, at most ten. From an established nonprofit with a programs section, at most a hundred. From a public body such as a city, a library or a school board, at most a hundred: we look for its pages about children, youth and families, not the whole site. We stop sooner when we find no more pages to look at. Sitemap files count toward these limits; robots.txt does not. If our crawler fails partway through a site, we may start that site once more. Separately, to check what we missed, we may once read a site's sitemap files in full (at most fifty) and up to ten of its other pages. Never more than a couple of hundred from one site.
- We come back occasionally to check whether something changed — at most monthly for details that change often, such as whether you are accepting new clients, and far less often for things like your address.
robots.txt
We follow the Robots Exclusion Protocol as specified in RFC 9309.
To block us specifically, add this to your robots.txt:
User-agent: AIWikiBot
Disallow: /
That takes effect on our next check, within 24 hours, and needs no email.
One detail worth stating, because implementations differ: if your robots.txt returns a server error, we treat it as "do not crawl" and stop — not as permission. A deploy that briefly breaks that file will pause us rather than unleash us.
How to make us stop
Any of these works, and none of them requires a reason:
- Add the
robots.txt lines above.
- Email crawler@educationall.tech. We will stop within two business days and add your site to a permanent exclusion list.
- Block our requests at your server or firewall. We will detect it, record it, and stop scheduling your site.
Opt-out is permanent. We do not re-approach a site that has asked us not to crawl it.
If your organization is listed
Our listings are built from published sources and may be wrong or out of date.
- To correct something, email crawler@educationall.tech with what is wrong. We will suppress the disputed detail while we check it, rather than leave it up while we deliberate.
- To have your listing removed entirely, email us and say so. We will remove it and keep a record that it was removed, so it does not reappear.
- To see what we hold about you, ask. If any of it is personal information about an identifiable individual, you have a right under PIPEDA to see it and to challenge its accuracy, and we will respond within 30 days.
You do not need an account with us to do any of this, and we will not ask you for one.
Accountability
A person is responsible for this crawler's behaviour. If it does something it should not — hits your site too hard, ignores your robots.txt, or records something it should not have — email crawler@educationall.tech and it will be stopped. We would rather hear from you than not.