Pangram Flags 25.7% of Long Social Posts as Fully AI-Generated

Pangram reported on July 9 that 25.72% of social posts longer than 250 words in its opt-in sample were flagged as fully AI-generated; LinkedIn exceeded 40%. The dataset covered 1,002,627 posts seen by users of Pangram's Chrome extension, so the findings are a useful contamination signal for social-text datasets, not a population estimate or ground truth for individual posts.
Pangram reported on July 9, 2026 that 25.72% of social posts longer than 250 words in its opt-in sample were flagged as fully AI-generated. LinkedIn had the highest longform rate at more than 40%, while X showed the largest combined share when fully generated and mixed AI-human articles were counted together.
The study offers a large observational signal about what participating users encountered in their feeds. It does not establish the share of all content published on any platform, and its labels come from Pangram's own detector.
What Pangram measured
Pangram said users who opted to share statistics from its Chrome extension contributed 1,002,627 unique posts from LinkedIn, Medium, Substack, X, and Reddit after the extension launched on April 24. The company included items longer than 50 words and treated items over 250 words as longform.
Across all scanned items, Pangram reported a 13.8% fully AI-generated rate. LinkedIn supplied about one-third of the sample but 62% of the content flagged as AI-generated. On X, 23.9% of articles were flagged as fully AI-generated and 22.9% as mixed or AI-assisted. The Register and TechRadar independently reported the study and its principal figures; TechRadar also obtained a response from LinkedIn saying it works to reduce low-quality, automated, or generic content.
Read the figures as detector output
The sample reflects feeds viewed by extension users who opted into research sharing, not a randomized census of each platform. Platform mix, user behavior, language, post length, and the detector's decision threshold can all affect the rates. Pangram states that its version 3.3 model has a 0.01% false-positive rate, but that performance claim is vendor-reported and does not independently validate every label in this social-media sample.
For data and ML teams, the practical result is a provenance warning. Social text used for training, evaluation, sentiment analysis, or market research may contain more synthetic material than its conversational tone suggests. Detector scores can support triage, but teams should combine them with sampling, source history, deduplication, temporal checks, and downstream robustness tests before treating authorship labels as fact.
Key Points
- 1Pangram flagged 25.72% of longform items in its opt-in social-feed sample as fully AI-generated, with LinkedIn above 40%.
- 2The 1,002,627-post dataset reflects content seen by participating extension users rather than a randomized estimate of all platform content.
- 3ML teams should treat detector labels as a corpus-risk signal and pair them with provenance checks, sampling, and downstream tests.
Scoring Rationale
The study quantifies a material data-quality signal across more than one million observed posts, but its opt-in sample and vendor-owned detector limit generalization and keep the impact measured.
Sources
Primary source and supporting public references used for this report.
Practice with real Social Media data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Social Media problems


