Skip to content
RedLens Get RedLens on GitHub
Public Discussion Analysis

How to Run an Ethical Online Background Check

A practical workflow for checking public online history without mistaking public access, identity matches or activity patterns for proof.

The RedLens Project

An ethical online background check answers a defined, relevant question with the least personal data necessary. It does not begin with “find everything.” Public availability is permission to view material, not a blank check to compile, identify and redistribute it.

That distinction matters on Reddit. Most communities are public, but public discussions can still contain sensitive or identifying information, involve people at risk, or reach an audience much larger than the author expected, as a 2024 systematic review of Reddit research ethics explains.

Set the rules before searching

Write a short case sheet before collecting anything:

  • Purpose: What specific decision or research question will this check inform?
  • Authority: Is this personal due diligence, journalism, academic research, trust-and-safety work or employment screening?
  • Scope: Which platforms, account identifiers and dates are relevant?
  • Exclusions: Which facts will you not collect or use—for example, health disclosures, family details or protected characteristics unrelated to the question?
  • Evidence threshold: What would confirm, weaken or leave the hypothesis unresolved?
  • Retention: When will the source material and working notes be reviewed or deleted?

If you cannot state a legitimate purpose and stopping rule, do not start. The Association of Internet Researchers treats ethics as an ongoing, context-specific process rather than a one-time checklist; it says research design should include “careful reflection” on context and the requirement to minimize risks and harms (AoIR Internet Research Ethics 3.0).

Check which legal regime applies

Ethics is not a substitute for legal compliance. Rules depend on the purpose, location, data source and person being checked.

In the United States, an employer using a company that compiles background information must follow the Fair Credit Reporting Act. Before obtaining the report, the employer must provide a stand-alone written notice and get written permission. Before taking adverse action based on it, the employer must provide the report and a summary of FCRA rights so the person can review and explain the information (FTC and EEOC guidance). State and local rules may add requirements.

Academic researchers should submit the protocol to their institution’s ethics or IRB process rather than assuming that public posts are exempt. Journalists, investigators and trust-and-safety teams should obtain review under their own legal and editorial policies when identification, allegations or vulnerable people are involved.

Use a reproducible, minimum-data workflow

1. Confirm the account before attributing it

A matching display name is not identity proof. Look for multiple independent, non-sensitive links: a self-linked website, consistent professional details, a cross-platform link posted by the account, or direct confirmation. Record the match as confirmed, probable, possible or unresolved, along with the basis for that label.

Reddit does not require an account holder to provide a real name, so a Reddit username should not be treated as a legal identity (Reddit Privacy Policy). Do not resolve ambiguity by collecting relatives’ details, breached data, private messages or deceptive access.

2. Collect only what answers the question

Use the narrowest relevant account list, communities, keywords and date range. For a Reddit check, archive only public material and keep the working dataset local where practical. If you need a reproducible starting point, the local Reddit history workflow explains how to preserve public posts and comments for analysis.

Pseudonymize usernames in analysis tables unless identity is necessary to the authorized task. Keep the re-identification key separate and access-controlled. Pseudonymization reduces casual exposure; it does not make distinctive quotations or posting histories anonymous.

3. Preserve context and provenance

For each item that may affect a finding, retain:

  • source URL and account identifier;
  • post or comment timestamp and collection time;
  • subreddit and parent-thread context;
  • exact search or query criteria;
  • whether the item was edited, removed or unavailable when reviewed;
  • a short note separating the observed fact from the analyst’s interpretation.

A post may be satire, quotation, role-play, a rebuttal or an outdated view. Read the surrounding thread and linked material before coding it.

4. Separate signals from conclusions

Write observations first: “Accounts A and B posted the same link within four minutes.” Then write the bounded inference: “This timing is consistent with coordination, but shared news alerts or community routines remain plausible.” Synchronized activity is a lead, not proof of shared control, payment or intent; see the coordination evidence standard.

Apply the same discipline to criminal-record material. An arrest does not establish that criminal conduct occurred, and the EEOC advises employment screens to consider the nature of the conduct, elapsed time and nature of the job, with individualized assessment where appropriate (EEOC guidance).

5. Corroborate consequential findings

Do not make a high-impact decision from one post, one name match, an automated sentiment score or an LLM summary. Check original records, chronology and plausible alternative explanations. Have a second reviewer examine both supporting and contradictory evidence without being told the preferred conclusion.

When a finding could materially harm someone, give that person a meaningful opportunity to correct identity errors, explain context and provide contrary evidence. Record corrections in the same case file rather than silently replacing the original conclusion.

6. Report less than you collected

Include only evidence needed to support the conclusion. Prefer aggregates or careful paraphrases when exact wording is unnecessary. Removing a username while reproducing a verbatim sentence may not protect the author because quoted public text can be found through search; research has documented this re-identification risk (Fiesler and Proferes).

Use calibrated labels such as observed, corroborated, inferred and unresolved. Avoid labels such as “dangerous,” “fake,” “bot” or “coordinated” unless you define the criterion and the evidence meets it.

Final review before a decision or publication

Confirm that:

  • the check stayed within its written purpose;
  • the account-to-person match is supported and uncertainty is visible;
  • protected or sensitive information did not influence an unrelated decision;
  • important claims have contextual and independent support;
  • the subject had a correction path when stakes warranted it;
  • quoted text and screenshots cannot cause avoidable re-identification;
  • access, retention and deletion dates are recorded;
  • another reviewer could reproduce the result from the case sheet without collecting extra personal data.

The defensible output is not the largest dossier. It is the smallest documented record that supports the stated conclusion—and makes its limits easy to inspect.