How it works

Thread Witness reads public discussion about a product, merges reports of the same problem or request into one topic, and counts the unique people behind each one. A language model writes the weekly report only from that stored evidence, and an automated verifier checks every citation, quote and number before release. Every factual sentence in the report links to a public post.

This page explains each step and the rules the report follows.

The pipeline

  1. Discover. Each source is read on a schedule through its official API, a published feed, or ordinary rate-limited page requests where its terms permit or do not prohibit automated reading. Each crawl re-reads the last 72 hours to catch late replies. In development: a daily search for new places where your product is discussed. A person reviews each new source’s terms before it is ever read.
  2. Pseudonymize. Author names are replaced with keyed pseudonyms, and email addresses, phone numbers and ID-like strings are removed, before posts enter our database or are sent to a language model. Usernames mentioned inside post text are removed from reports by the verifier.
  3. Match to product. Each post is tagged with the products it is about. One thread can be about several products.
  4. Classify. Each post is typed: a first-person issue report, a feature request, a workaround, a confirmation or refutation of a workaround, an official response relayed secondhand, praise, or other.
  5. Extract. Structured fields are pulled from each relevant post: symptom, error codes and on-screen text, triggering conditions, version, environment, the workaround tried and whether it worked, the capability requested and why. Every extracted value must point to the exact words in the post. A value without them is rejected.
  6. Merge. Reports of the same problem become one issue; requests for the same capability become one request.
  7. Score. Counts and scores are computed from the stored data.
  8. Write and verify. A language model writes each section only from the evidence for that topic, under an instruction to omit rather than guess. The verifier then checks the draft.
  9. Deliver. The report, a PDF, an evidence file and a one-page executive summary are produced. Alerts and ticket sync are built and being prepared for production.

These steps are built. Reading public sources runs today; the later steps are in testing during early access.

Topic IDs

Every issue and feature request gets an ID such as NWC-ISS-4H8T2C or NWC-FR-6D1K8V. The ID never changes and is never reused, so you can put it in a ticket, a knowledge-base article or a review deck and find the same topic next quarter. When two topics turn out to be one, or one turns out to be two, the merge or split is logged and listed in the report.

Issues and requests are separate

Problems and wishes are extracted and scored separately. Priority and Demand are both scored 0 to 100.

A topic enters a main table only when at least 3 unique first-person reporters described it in the trailing 8 weeks, unless it is in the safety or security class. Below that it appears in an “Emerging (low evidence)” table.

Severity and human review

Each first-person report is classed by what the reporter experienced:

Class Meaning
S1 Safety or security Could plausibly cause injury, loss of control, a security breach, exposure of protected data, or disable a safety or security function.
S2 Data loss, lockout or unavailable The product is unusable, data is irrecoverable, or recovery needs vendor support or a reset that loses data.
S3 Core function loss A primary function fails, but the product is otherwise usable or recovers with a user action.
S4 Degraded Works with reduced quality, intermittent failures or slower performance.
S5 Cosmetic Presentation, wording or minor annoyance.

A topic’s severity is the highest class supported by at least two independent reporters. Anything in the safety or security class is shown as “pending human review” until a person has reviewed it and recorded a decision.

Workaround counts

Users often post a workaround and others reply that it did or did not work. Thread Witness counts those replies:

Labels follow fixed rules: Confirmed (N), Confirmed, temporary (N), Likely (N), Mixed, Partial, Unconfirmed, Refuted, Secondhand. These label names count user replies; they are not our conclusions. Each workaround also carries its cost and risk: requires admin, requires a vendor ticket, disables a security function, risks data loss, needs third-party software, time burden. The report does not endorse workarounds; it reports what users said.

What users praise (in development)

Posts that say what people like are already recognized when posts are classified. In development: grouping them into strength topics and showing them beside issues and requests, so a team sees what to protect as well as what to fix.

Release correlation

Release notes, changelogs and status-page feeds are read as events. Each issue timeline is annotated with releases and incidents, so a rise in reports can be read against what shipped or what went down that week. The report describes timing; it does not claim cause.

Fix verification (Business and Enterprise plans)

When you tell us an issue was fixed in a release (in the dashboard, or by email during early access), Thread Witness compares public reports before and after the fix ships, from week 4 onward. Verdicts describe public reports, not product behavior.

Verdict In plain language
Fix verified Reports fell to a quarter of the pre-fix rate or less for two weeks, most reporters are on the fixed version, and almost none of them still describe the problem.
Partially effective Reports fell but not far enough, or a few people on the fixed version still describe it.
Inconclusive Too early, too few people on the fixed version yet, or too little volume before the fix to compare.
Regression suspected Several people on the fixed version describe the problem, or reports stayed high or rose.
New regression Reports of a different problem that start at the fixed version.

Fix verification needs a baseline and several weeks of observation, so it runs on Business and Enterprise subscriptions.

Claim types

Every factual sentence in a report carries one of four badges:

Badge What it means How it is checked
Reported A statement the cited posts make directly. Checked for support against each cited post.
Aggregated A number computed from the stored data. Recomputed before release.
Inferred Our interpretation, hedged and pointing to its evidence. Checked for hedged wording.
Secondhand A user relaying what someone else said. Shown as relayed; never counted toward severity.

You will see phrasing such as “N reporters describe”, “community reports indicate”, “consistent with” and “suggests investigating”. No individual, reseller, dealer or third-party vendor is named as being at fault.

The verifier, in plain language

Before a report is released, nine checks run on every sentence:

  1. Every citation points to a stored public post in the report’s window.
  2. Every quote matches the cited post exactly.
  3. Every number recomputes from the stored data.
  4. Every error code, alert text or version string appears word for word in a cited post.
  5. Every Reported statement is supported by the posts it cites.
  6. Every sentence that states a fact has a citation.
  7. No usernames, personal data or full posts appear anywhere.
  8. Our interpretations use hedged wording, and conclusion words such as “proves”, “confirms” or “caused by” appear only inside quotes. Fixed label names such as Confirmed (N) count user replies and are the one exception.
  9. Every safety- or security-class topic has a human review record.

A sentence that fails is removed, rewritten, or downgraded from Reported to Inferred. A report that fails check 7 is not released. A safety-class topic without a review record is held as pending human review and stays out of the report’s reviewed safety findings. The report header prints the tally.

What the counts mean

Counts measure public discussion in the sources we read. They are not failure rates, incidence rates or satisfaction scores. People who post in public are not a random sample of your users: some groups post more than others, some problems are more likely to be discussed, and the sources we can read differ by product. Every report lists its sources, the date range, what was scanned, and the known gaps. Thread Witness publishes no scorecards and no rankings of vendors.

See a sample reportHow we source data