How we source data
Thread Witness reads only public information, and only where we have a recorded basis to read it.
Typical sources
- Public issue trackers, such as GitHub issues and public Jira projects
- Stack Exchange sites (counted and linked, not quoted)
- Hacker News
- Vendor developer forums and public help-center communities
- Release notes, changelogs and status pages
- For consumer brands and home builders (in development): CPSC SaferProducts.gov incident reports and recalls, and NHTSA vehicle complaints and recalls, through their official open-data interfaces
If you run a community whose terms restrict automated reading, written permission from you adds it as a source for your reports.
Official APIs and feeds first
We prefer official APIs, published RSS and Atom feeds, public status pages and government open data. Where a site offers none of these, we make ordinary rate-limited page requests, and only if its terms permit or do not prohibit automated reading and its robots.txt allows our agent.
A recorded basis for every source
Each source carries a record of the basis on which we read it: reviewed API terms, government open data, site terms that permit or do not prohibit automated reading, a published feed, written permission from the operator, or a license. Reviews are renewed at least every 180 days.
Where a site’s terms restrict automated or commercial use but the site publishes an RSS or Atom feed, we read only that feed, not the site’s pages. Sources we do not read are listed as known gaps in every report.
How we read
- robots.txt is checked before every crawl.
- Our crawler identifies itself with the user agent
ThreadWitness/1.0 (+https://threadwitness.com/data-sources; hello@threadwitness.com). To block it, disallow the robots.txt tokenThreadWitness. - No logins, no CAPTCHA solving, no proxy rotation, no circumventing paywalls or bot protection. If a site blocks us, we stop.
- Requests are rate-limited per site.
What we do not read
- Review sites and retailer reviews.
- Reddit, except under a signed license. We do not hold one today.
- Anything behind a login, private groups, direct messages or email.
Authors are pseudonymized
Author names are replaced with keyed pseudonyms before posts enter our database or are sent to a language model. Email addresses, phone numbers and ID-like strings are removed at the same step. Usernames that appear inside post text, such as a reply that mentions another user, are removed from every report by the verifier before release. We build no profiles of individuals: authors exist only as pseudonyms attached to counts.
Raw copies of fetched pages, which can include usernames, are held in a restricted cache so posts can be reprocessed. They are not used in reports. Automatic deletion of those copies after a set period is in development.
Quotes
Quotes are verbatim, at most 200 characters, linked to the source post, and never a full post. Each person is quoted at most once per topic. Some sources, such as Stack Exchange, are counted and linked but not quoted.
Edits and deletions
Each crawl re-reads the most recent 72 hours of a source to catch late replies. In development: updating stored posts that were edited at the source, and a monthly re-check of posts cited in stored reports that drops quotes from posts deleted or edited since. Copies of reports already delivered into a customer’s own systems cannot be recalled by us.
AI providers
To classify posts and write reports we will use AI model and embedding providers. Before any collected data is sent to them, this page will name each provider, what it receives and its retention terms. They will receive post text with author names replaced by pseudonyms. We do not train models on collected content.
If you posted in a source we read
- Who we are. Thread Witness. Contact: hello@threadwitness.com.
- Purpose. We count and cite public reports about products so that the teams responsible for them can find and fix problems.
- Lawful basis. Our legitimate interest in analysing public product discussion, limited by the safeguards on this page: public sources only, pseudonyms instead of usernames, short linked quotes and no profiles.
- What we hold. The public post text, its link and date, and a pseudonym in place of your username. Raw page copies are described above.
- Retention. Our retention period for stored post text is 24 months. Automatic deletion at the end of that period is in development.
- Your rights. You can ask what we hold about your posts, object to our use of them, and ask us to correct or erase them. Write to us as described on Corrections and takedowns. We may ask you to show that you control the account before we act.
- Complaints. If you are in the UK or the European Economic Area, you can complain to your data protection supervisory authority.
Corrections and takedowns
Anyone can ask us to correct or remove something. We acknowledge every request within 48 hours. How to request a correction