Skip to content

Inside the Engine

Nobody has ever mapped the digital infrastructure of the entire American nonprofit sector. So we built the machine that does.

This page is for the people who want to look under the hood: the funders deciding whether we’re worth backing, the researchers wondering if our data is real, and the nonprofit leaders asking “wait, you scanned my website?”

Yes. Here’s how it all works, and why it matters.

The Problem Nobody Was Solving

Every US nonprofit files paperwork with the IRS. That paperwork is public. Buried inside it is the answer to a question nobody had ever seriously asked:

What does the digital infrastructure of an entire charitable sector actually look like?

Not a survey of 400 organizations that had time to answer a questionnaire. Not a vendor report designed to sell you something. The real thing: every organization, every website, measured directly.

The data existed. The engineering to connect it didn’t. IRS records don’t come with a “website” field you can trust. Nonprofit sites go up, go down, move, and get abandoned without anyone updating a registry. Connecting a legal entity in a tax filing to a live server on the internet, times two and a half million, is a hard engineering problem.

That’s the problem we solved. Solving it is what makes this dataset exist.

What the Engine Actually Does

TSNI Nonprofit Web Intelligence is a pipeline. Public records go in one end; sector-wide intelligence comes out the other. Four stages.

1. Map

We start with the IRS’s public files and build a complete registry of America’s tax-exempt organizations: roughly 2.5 million legal entities, enriched with three-quarters of a million full tax filings. Identity, location, mission category, finances, and history, all in one queryable warehouse.

Then comes the hard part: matching each organization to its real website. Automated pipelines, AI research agents, and human volunteers work through the registry together. So far, roughly 700,000 organizations have a validated live website.

Which leaves the number that should stop you cold: the majority of US nonprofits have no findable website at all. Not a bad website. None. We believe we are the first to measure the true size of that gap.

2. Watch

Our crawler visits the sector’s websites the polite way: read-only, rate-limited, honoring every robots.txt, always identifying itself. It archives what the public sees, from homepages down to the pages that matter most: about, programs, donate.

We’ve archived and analyzed hundreds of thousands of pages to date, and every crawl cycle adds history. Snapshots become trend lines. Trend lines become the sector’s first longitudinal record of its own digital health.

3. Understand

Raw HTML is noise. The engine reads every archived page and fingerprints what’s actually running on it: which CMS, which donation platform, which analytics tools, which security posture. Our detection catalog covers 363 distinct technologies and grows continuously, because the engine also discovers tools it has never seen before and flags them for review.

This is how we know, with receipts, which platforms dominate the sector, which ones are abandoned relics, and which security problems are quietly widespread.

4. Rank

Finally, everything feeds the TSNI Nonprofit Web Index: our composite score of real digital effectiveness, combining measured traffic authority, performance, and security into a single comparable number. Refreshed weekly, tracked over time. It powers our Top 100 lists, state-by-state comparisons, and category rankings.

When Alexa died, the world lost its ranking of nonprofit websites. Ours is better, because it’s built only for this sector, on measured data.

The Map Nobody Else Has

There’s a whole category of technology that exists only because nonprofits exist. Nobody has ever catalogued it. We are.

The Blind Spot

There are tools out there that can look at any website and tell you what it’s built with. They’re genuinely good at it, if you run an online store. They’ll spot a retailer’s checkout system, a startup’s analytics, a media company’s ad network in seconds. That entire industry was built to map the technology of commerce, because commerce is where the money is.

So here is what those tools cannot see. The donor-management system your local shelter depends on. The volunteer-scheduling platform running a disaster-relief network. The peer-to-peer fundraising engine behind a hospital gala. The church-management software running a congregation. The school-website platform a scholarship fund is built on.

These are not fringe tools. They are the operational backbone of the charitable sector. Hundreds of distinct platforms that millions of organizations rely on every single day. And to the rest of the technology world, they are invisible, because the sector was never a lucrative enough market for anyone to bother mapping it.

That blind spot is the opportunity. We built the engine that sees into it.

The Catalog That Grows Itself

This is the part we’re proudest of. It’s also the part that makes this defensible.

A fixed catalog would go stale the day we shipped it. Ours doesn’t, because the engine is built to find what it doesn’t already know. When it runs into technology it has never seen before, some tool quietly powering a corner of the sector, it doesn’t shrug and move on. It flags the unknown. Then our AI research agents identify it, verify it, and fold it into the catalog.

So the map fills itself in. Every crawl cycle, the engine gets a little less blind. Tools that no detection service on earth has a name for get discovered, identified, and added. And once they’re in, we can see them everywhere they show up across the sector.

That’s the flywheel. The more we watch, the more we discover. The more we discover, the sharper the map gets. We don’t publish how the discovery works, and we never will, because a map this valuable only stays trustworthy if it’s hard to game. But you can see the result plainly enough: a catalog that started with the obvious and now reaches software the commercial world doesn’t even have a name for.

We developed a completely unique dataset on the technological life of an entire nonprofit sector. Owned by a nonprofit. Given back to the sector. Independent and completely free from corporate influence.

How Two Engineers Run All of This

Fair question. Organizations with 50x our headcount run smaller data operations.

The honest answer: we treat AI agents as staff. Our engine delegates the grindwork (verifying websites, triaging unknown technology signals, researching organizations that automated matching can’t crack) to AI agents working through supervised queues. Every agent decision lands in a review ledger. Humans set policy, review output, and take responsibility; agents provide the scale.

We think this is what every high-leverage nonprofit will look like in ten years. We’re just early.

What we won’t publish: our matching methodology, data source contracts, or detection signatures. Not because we’re precious about it, but because this dataset only stays trustworthy if it’s hard to game. If you’re a researcher with a legitimate need, talk to us.

The Rules We Scan By

Scale without ethics is just surveillance. Our rules are absolute, and they are the same five rules we publish on What We Do: passive only, fully identified, boundaries respected, organizations rather than people, and never weaponized.

Any organization can opt out with one email, no questions asked.

What This Unlocks

If you run a nonprofit: a free, objective outside view of your digital health, in plain mission language, not engineer-speak. No sales call afterward, because we have nothing to sell you.

If you’re a researcher or sector body: the first complete, longitudinal, measured (not surveyed) dataset on nonprofit technology. We’re actively looking for research partners. Our upcoming census-data integration will map digital health against community demographics, down to the communities being left behind.

If you’re a funder: this is infrastructure the sector has never had, built and operated at a fraction of what it should cost, currently paid for out of two engineers’ pockets. Your support doesn’t fund overhead. It funds compute, data, and the roadmap on What’s Next. We’ll show you the ledger.

Partner with us → · Pitch in → · Back to the overview →


Note: TSNI Nonprofit Web Intelligence is an analytical and monitoring platform. TSNI does not provide IT support, technical consulting, or website remediation services based on these findings.