How we collect, verify, and decode every conversation on X
A transparent look at the data pipeline, statistical safeguards, and editorial review behind every TweetBlocker insight — from the first API call to the final analyst note.
Overview
Built on forensic data principles, not black-box scraping
TweetBlocker treats every public post on X as a primary source document. Our methodology blends deterministic API ingestion, multi-stage verification, and human editorial review so that the narratives you read on our platform are traceable, reproducible, and properly contextualized.
We publish this page for three reasons: to let researchers audit our work, to help partners integrate our data with confidence, and to set honest expectations about what our platform can — and cannot — tell you about activity on X.
01 · Data Pipeline
Five stages from raw post to verified insight
Every datapoint on TweetBlocker passes through the same deterministic pipeline. Each stage logs an immutable audit record so analysts can trace any conclusion back to its source.
Ingest
Authenticated X API v2 requests scoped to public posts, filtered by topic, account, or time window.
Normalize
Unified schema: UTC timestamps, resolved entities, canonical URLs, language detection on 100+ locales.
Enrich
Thread reconstruction, quoted-post linking, media fingerprinting, and account-level metadata joins.
Verify
Statistical deduplication, bot-likelihood scoring, and human spot-checks against original posts.
Publish
Versioned release to the dashboard, partner API, and weekly Insights digest with full provenance.
02 · Sources & Coverage
What we collect — and the honest limits of that collection
TweetBlocker relies on official, authenticated X API endpoints as its primary source. We supplement with publicly cached snapshots and partner feeds, but we never ingest private messages, suspended-account content, or posts removed under legal process.
Primary: X API v2
Authenticated access to the Search, Timelines, and Lookup endpoints. Requests are signed, rate-limited per our plan, and logged for audit. We respect all `next_token` cursors and never paginate past the 7-day historical window unless explicitly extended by X.
- Public posts only — no DMs, no protected accounts
- Rate-aware backoff with exponential jitter
- Full response payload stored in cold storage for 18 months
Secondary: Cached Snapshots
For events where X search returns gaps — breaking news, mass suspensions, regional throttling — we lean on third-party archives such as the Internet Archive's X collection and academic partnerships. These snapshots are flagged with a "cached" provenance marker in every UI surface.
- Cross-referenced against at least one live API query when possible
- Capped at 15% of any single Insights report
- Never used as a sole source for a headline metric
Account Metadata
We join every post with the author's public profile state at the time of the post — display name, bio, follower count, verification status, and account creation date. This is critical for cohort analysis and for catching coordinated networks that rebrand.
- Snapshot per post, not per query window
- Profile changes are diffed and logged
- Suspended accounts are retained with a `was_suspended` flag
03 · Verification & Accuracy
We measure accuracy in public, and correct in public too
Every model on TweetBlocker ships with a published confidence interval, a known error budget, and a correction log. If we miss, you'll find it — and the fix — on our Insights page within 72 hours.
Honest limits we publish with every chart
API search returns public posts that match a query, not the full universe of posts about a topic. We never infer what we did not see, and our reach metrics are scoped to the post sample retrieved, not total platform activity. When X changes its API surface, we flag affected historical comparisons in the methodology footer of each Insights piece.
04 · Ethics & Privacy
Public data, treated like private data
TweetBlocker only ingests posts that are publicly visible on X without authentication. We treat that data with the same care we'd want for our own — minimizing re-identification, separating content from identifiers when possible, and giving account holders a clear path to opt out of amplification on our platform.
No PII enrichment
We do not link X handles to emails, phone numbers, physical addresses, or any third-party identity graph. Account-level features are derived solely from public profile metadata at the time of posting.
Respect for takedowns
When a post is deleted on X, we remove it from active surfaces within 24 hours. Historical reports that cited the post are annotated with a `since-removed` note rather than silently rewritten.
Opt-out & corrections
Any account holder can request exclusion from TweetBlocker's surfaced analysis. We honor verified requests within 5 business days and publish anonymized quarterly transparency stats on the volume of opt-outs processed.
05 · Frequently Asked
Questions we hear most often from partners and press
How current is the data I see in the dashboard?
Default dashboards refresh every 15 minutes for active topics and every 6 hours for long-tail queries. The "last fetched" timestamp is exposed on every chart so you always know the freshness of the underlying sample.
Do you include deleted, edited, or suspended posts?
Deleted posts are kept in cold storage with a `since_removed` flag and excluded from active charts. Edited posts are tracked as versions, and we surface the most recent version in dashboards with an "edit history" affordance. Suspended posts remain in historical analyses with a `was_suspended` marker, so the original signal is never silently erased.
How do you detect coordinated or inauthentic activity?
We score accounts on a 0–100 inauthenticity index using temporal posting patterns, semantic similarity of recent posts, follower/following graph structure, and shared media fingerprints. Accounts above 70 are surfaced in a separate "review queue" tab and are never used as a single source of attribution in a headline insight.
Can I audit a specific metric you published?
Yes. Every public chart links to a methodology card with the query, time window, sample size, and a reproducible notebook for paid research partners. Send a request via the Contact page and we'll provision access within two business days.
What are you explicitly not trying to measure?
We don't claim to measure total platform activity, sentiment in a clinical sense, or the off-platform impact of any single post. TweetBlocker is a forensic X-data platform — not a general social listening tool — and we keep the scope of our claims narrow on purpose.
The Editorial Team
Analysts behind the methodology
TweetBlocker's methodology is maintained by a small, named team of data engineers, statisticians, and former investigative reporters. Every correction and model update is signed.
Dara Reyes
Former data editor at a major investigative newsroom. Designs the verification protocol and signs off on every model release.
Marcus Khoury
Owns the ingest and normalization layer. Built the audit log format that underpins our reproducibility guarantees.
Saoirse Ní Bhroin
PhD in computational social science. Reviews confidence intervals and runs the quarterly accuracy benchmarks.
Jonas Tellez
Reviews every dataset against our ethics charter and handles opt-out and takedown requests from account holders.
See the methodology in action
Browse a live analysis built on this exact pipeline — every chart links back to the query, the sample, and the audit log.
Read the Latest Insights Compare Plans