How we source data

Last updated: 11 October 2026

Thread Witness reads only public information, and only where we have a recorded basis to read it.

Typical sources

If you run a community whose terms restrict automated reading, written permission from you adds it as a source for your reports.

Official APIs and feeds first

We prefer official APIs, published RSS and Atom feeds, public status pages and government open data. Where a site offers none of these, we make ordinary rate-limited page requests, and only if its terms permit or do not prohibit automated reading and its robots.txt allows our agent.

A recorded basis for every source

Each source carries a record of the basis on which we read it: reviewed API terms, government open data, site terms that permit or do not prohibit automated reading, a published feed, written permission from the operator, or a license. Reviews are renewed at least every 180 days.

Where a site’s terms restrict automated or commercial use but the site publishes an RSS or Atom feed, we read only that feed, not the site’s pages. Sources we do not read are listed as known gaps in every report.

How we read

What we do not read

Authors are pseudonymized

Author names are replaced with keyed pseudonyms before posts enter our database or are sent to a language model. Email addresses, phone numbers and ID-like strings are removed at the same step. Usernames that appear inside post text, such as a reply that mentions another user, are removed from every report by the verifier before release. We build no profiles of individuals: authors exist only as pseudonyms attached to counts.

Raw copies of fetched pages, which can include usernames, are held in a restricted cache so posts can be reprocessed. They are not used in reports. Automatic deletion of those copies after a set period is in development.

Quotes

Quotes are verbatim, at most 200 characters, linked to the source post, and never a full post. Each person is quoted at most once per topic. Some sources, such as Stack Exchange, are counted and linked but not quoted.

Edits and deletions

Each crawl re-reads the most recent 72 hours of a source to catch late replies. In development: updating stored posts that were edited at the source, and a monthly re-check of posts cited in stored reports that drops quotes from posts deleted or edited since. Copies of reports already delivered into a customer’s own systems cannot be recalled by us.

AI providers

To classify posts and write reports we will use AI model and embedding providers. Before any collected data is sent to them, this page will name each provider, what it receives and its retention terms. They will receive post text with author names replaced by pseudonyms. We do not train models on collected content.

If you posted in a source we read

Corrections and takedowns

Anyone can ask us to correct or remove something. We acknowledge every request within 48 hours. How to request a correction