What Reddit’s Scraping Ruling Actually Covers
Judge Engelmayer’s ruling targets alleged CAPTCHA circumvention and deletion failures—not every local archive of public Reddit discussions and posts.

No. The July 2026 SDNY ruling did not hold that scraping public Reddit posts is illegal. It allowed two claim theories to proceed based on allegations that SerpApi bypassed Google’s CAPTCHA-based SearchGuard system and that third-party copies frustrated Reddit’s handling of user deletions. A personal local archive using a permitted access path and propagating deletions falls outside those allegations, although that does not create a safe harbor from contract, copyright, privacy, or other law (MediaPost).
The decision was procedural. Judge Paul Engelmayer did not find Perplexity or SerpApi liable, and the alleged conduct remained unproven at that stage. The ruling says that Reddit pleaded enough for specified claims to continue, not that collecting publicly visible Reddit content is categorically unlawful.
This is information about the reported court record, not individualized legal advice. A legal assessment still depends on the jurisdiction, access method, governing terms, material collected, retention policy, and intended use.
Why the Broad-Risk Reading Sounds Plausible
The consensus reading is understandable. “Reddit’s scraping lawsuit survives dismissal” sounds like a court has accepted the proposition that scraping Reddit is itself actionable. Reddit also has platform rules restricting how its data may be accessed and used, while public posts can contain copyrighted expression, personal information, and sensitive account histories.
Reddit’s displayed Data API Terms require compliance with its developer rules, documentation, approved use case, identification requirements, and technical limits. They prohibit circumventing API restrictions and describe the API license as revocable, non-transferable, and non-sublicensable. They also address deletion of content that is no longer needed and separate authorization for uses outside the approved scope (Reddit Data API Terms).
That makes “public” an incomplete answer. A page loading without a login does not automatically authorize indefinite storage, republication, resale, account-level profiling, or AI training. A project can avoid the conduct alleged in this case and still face a contractual, copyright, privacy, or platform-enforcement dispute.
The consensus is therefore right about one point: public visibility is not a blanket exemption. It goes too far only when it treats this procedural ruling as a holding against public-content collection generally.
Compare Your Archive With the Alleged Conduct
Set the three answers to match your workflow; the result identifies which parts of the ruling’s alleged pattern overlap.
Ruling Comparison
Does Your Setup Match The Alleged Conduct?
Choose the facts of your workflow. This compares conduct only with the allegations described in the reported ruling; it does not decide whether the workflow is lawful.
How Each Answer Maps To The Court Record
The two surviving theories and the separate commercial distinction should not be collapsed into one general “scraping” rule.
| Workflow Fact | Reported Case Connection | What It Changes | What It Does Not Decide |
|---|---|---|---|
| No bypass; no resale; honors deletions | Outside alleged pattern | Removes the two direct factual matches and the commercial resemblance | Terms, copyright, privacy, provenance, or jurisdiction |
| Bypasses a CAPTCHA or access control | DMCA allegation | Matches the reported basis for the anti-circumvention theory | Liability; the ruling was procedural |
| Fails to propagate deletions | Privacy allegation | Matches the alleged interference with user-deletion commitments | Whether harm occurred; no final finding was made |
| Resells or licenses collected records | Commercial resemblance | Makes the workflow more like the reported SerpApi supply setting | Resale was not itself the stated anti-circumvention holding |
Scope: “Outside the alleged pattern” is not a safe harbor. Access authorization, Reddit’s terms, copyright, privacy, source provenance, and local law remain separate questions.
Sources: MediaPost’s report on Reddit v. Perplexity and SerpApi; Android Headlines’ report on Google v. SerpApi. Allegations were not final findings.
The visible default assumes no bypass, no resale, and deletion requests honored.
The checklist is deliberately narrow. “Outside the allegations” does not mean “legally approved.” It means the three selected facts do not match the conduct on which the reported anti-circumvention and deletion theories depended.
There is also no numerical request threshold in the reviewed reporting that separates lawful from unlawful collection. The relevant distinction in this ruling was the alleged defeat of a particular technical control, not whether a collector made an unspecified number of requests.
The Surviving Claims Turn On Circumvention And Deletions
Reddit alleged that automated tools obtained Reddit material from Google search results by bypassing Google’s CAPTCHA-based SearchGuard system. Judge Engelmayer allowed Reddit’s principal Digital Millennium Copyright Act anti-circumvention theory to continue because those allegations, if proved, could support the claim (MediaPost).
That is narrower than holding that copying a public post violates the DMCA. The reported theory depends on an alleged technological measure and deliberate bypass. Visiting a page available to an ordinary unauthenticated browser is not the same alleged conduct as automating the defeat of a CAPTCHA-based restriction.
The privacy-related theory was also tied to deletion. Reddit alleged that copies retained by third parties interfered with its ability to honor commitments to users who deleted content. The judge treated that alleged interference as capable of supporting Reddit’s claimed reputational injury at the pleading stage. He did not make a final finding that the defendants caused that harm.
For a local archive, deletion behavior therefore matters. An archive that detects removals and deletes corresponding local records differs from a system that intentionally preserves and distributes material after the user removes it. Cached copies, search indexes, exports, backups, reports, and derived tables can all retain the same information, so deleting only the primary database record may not complete the process.
The Commercial Pipeline Is A Material Difference
The dispute also concerns commercial data supply rather than solely personal archiving. Reporting on the parallel Google v. SerpApi litigation named Nvidia, Uber, and Adobe among SerpApi’s clients. That places SerpApi’s activity in a commercial service context distinct from a person keeping a local research corpus that is not licensed or sold (Android Headlines).
Commercial status alone was not the anti-circumvention holding, and personal use is not an automatic exemption. The distinction matters because resale and customer access add contractual, redistribution, provenance, and deletion questions that do not arise in the same way when records remain on one researcher’s machine.
The parallel Google case also shows why the surviving Reddit claims should not be mistaken for a final endorsement of every theory. Google’s case against SerpApi was dismissed on standing grounds, with a 21-day window to amend. The reporting characterized the standing problem as the platform not automatically owning the public search-result content at issue and said Reddit could encounter similar standing questions as its case proceeds (Android Headlines).
A dismissal on standing grounds does not establish that SerpApi’s collection was lawful. Likewise, denial of a motion to dismiss does not establish that Reddit will ultimately prevail. Both decisions concern whether particular parties and pleaded theories can proceed.
Local Storage Changes Exposure, Not Authorization
Keeping an archive locally can reduce distribution and security exposure. It does not answer how the records were acquired, whether complete posts can be copied, or whether deleted material can be retained.
A local archive is furthest from the reported allegations when it combines three characteristics:
- It does not defeat a login, CAPTCHA, block, rate restriction, or other technical protection.
- It is not sold, licensed, or exposed as a customer-facing full-text data service.
- It detects deletion signals and removes affected material from the places where it was copied.
The first and third points map directly to the reported claims. The second distinguishes personal archiving from the commercial pipeline described in the related reporting, but it is not an independent judicial safe harbor.
The acquisition route still needs scrutiny. Using an approved API within its current terms generally presents fewer authorization questions than borrowing credentials, entering a restricted community, solving CAPTCHAs automatically, or rotating identities to evade a block. An approved API does not resolve downstream copyright, privacy, retention, publication, commercial-use, or AI-training questions.
Public browser access also does not settle those downstream issues. General web-scraping guidance distinguishes access to public, unauthenticated pages from entry through technical barriers while cautioning that contract, copyright, and privacy rules can apply independently (Apify).
Full Posts And Account Histories Carry Separate Risks
A date, vote count, or subreddit name differs from the original wording of a post, photograph, illustration, or video. Collecting only factual fields needed for aggregate analysis can reduce the amount of expressive and user-linked material retained. It does not make all metadata unrestricted.
A username combined with timestamps, subreddit participation, location clues, and inferred beliefs can form a persistent behavioral history. Replacing the username with a stable pseudonym still allows every record carrying that replacement identifier to be linked. Full text may also reveal names, workplaces, locations, health information, or other identifying details.
The downstream design matters as much as storage location. A private corpus used to produce aggregate counts differs from a searchable archive displaying complete posts. A customer download containing raw text differs from a report of community-level trends. A supplier’s statement that records are “public” does not establish that it had authority to collect, resell, or license every intended use.
Copyright presents the same separation between access and reuse. Retrieving a post does not by itself grant permission to reproduce, distribute, publicly display, sell, or train a model on it. Reddit may host the material while the user or another rightsholder owns relevant rights. Permission from one party may not resolve the rights controlled by another.
RedLens Users Must Check The Actual Source Path
RedLens says it retrieves public Reddit discussions through the Arctic Shift mirror, stores them in a local SQLite file without requiring a Reddit account or API key, and supports exports and username pseudonymization (RedLens workflow). Those features describe the product’s operation; they do not establish the mirror’s authority, downstream rights, or deletion compliance.
Because that workflow uses a third-party archive, it should not be described as collection through Reddit’s own access path. A mirror changes where the user gets the data, but it does not eliminate questions about how the supplier originally collected it, which terms applied, whether removals are communicated, or whether the supplier may provide the records for the user’s purpose.
For a RedLens archive, the court record points to two concrete checks. First, determine whether any upstream collector bypassed a CAPTCHA or another technical control. Local operation does not erase upstream provenance. Second, determine whether the source communicates removals and whether those signals reach the SQLite database, exports, indexes, backups, and reports.
If the provenance is unknown, mark it unknown rather than assuming that public visibility resolved authorization. If the feed cannot communicate deletions, the archive cannot claim to match the deletion-respecting default in the checklist.
The Practical Boundary From This Ruling
The ruling supports a specific boundary, not a universal answer. Deliberately defeating a CAPTCHA-based control can support an anti-circumvention theory at the pleading stage. Preventing a platform from carrying out user deletions can support an alleged privacy-related injury. Neither holding says that copying any publicly visible Reddit post is illegal.
A personal archive with no technical bypass, no commercial redistribution, and working deletion propagation is outside the conduct alleged in this case. Its remaining questions concern the actual access authorization, Reddit’s applicable terms, rights in copied material, privacy consequences of compiling account histories, and the rules of the relevant jurisdiction.
If collection depends on defeating a login, CAPTCHA, block, or rate restriction, stop before engineering around it. If the project republishes full content, sells access, creates persistent user dossiers, preserves removed material, or trains an AI model, the narrow comfort supplied by this ruling does not answer the project’s larger legal questions.