The full research file on Reddit powermods, government-platform coordination, the censorship infrastructure network, AI moderation failures, and the Dead Internet.
Contents 54 sections

Research file for Book 3 (Lurk More), Chapters 7, 19, and 20. Also feeds thefire.lol reference page and bibliography.

All claims sourced. DEFAMATION NOTE headers on all named-person claims. Last updated: 2026-02-28


1. Reddit Powermod Networks

The 5/92 Stat

In May 2020, data visualizations went viral showing that 5 moderators controlled 92 of the top 500 subreddits. Reddit users who shared or discussed the data were banned. A broader analysis showed 6 moderators controlled over 118 of the top 500 subreddits.

DEFAMATION NOTE: The banning of users who shared the data is documented in community archives. Reddit made no public statement contesting the data.

Named Powermods

GallowBoob (Robert Allam)

Formerly Reddit’s most prolific poster, moderating 113+ subreddits. Was hired by UNILAD, later Jukin Media and BoredPanda, based directly on his Reddit influence. Accused of paid promotions — notably a Netflix post that appeared suspiciously early. Left Reddit in 2020 after being doxxed. Later hired as Director of Social Media at Tempo Storm.

DEFAMATION NOTE: GallowBoob’s real name, employment history, and departure are documented in public interviews and media coverage. The Netflix paid-promotion accusation is community-documented, not confirmed by GallowBoob or Netflix.

awkwardtheturtle

Moderated “hundreds, if not thousands” of subreddits. Documented behavior: stealing users’ posts, banning the originals, re-uploading as their own. A Change.org petition was launched against them. Eventually permanently suspended from Reddit after sustained community campaigns.

Other known powermods: IranianGenius, Cyxie, Merari01, Siouxsie_siousv2

The Default Subreddit System (2006–2017)

When users created a Reddit account, they were automatically subscribed to a curated list of “default subreddits” — r/pics, r/funny, r/videos, r/news, r/gaming, etc. Whoever moderated those defaults was effectively a gatekeeper for what new users saw first. The system was eliminated May 31, 2017 and replaced with r/popular and a personalized Home feed.

The Moderator Cap (December 2025)

In November 2025, a moderator on r/Art banned artist Hayden Clay Williams for mentioning “prints” (interpreted as self-promotion), deleted his post history, permanently banned him, and blocked appeals. When users protested, the head moderator went rogue, removed all other moderators, and locked the subreddit. This directly triggered Reddit’s new policy: maximum 5 communities with over 100k visitors per moderator. Enforcement began December 2025.

The Aimee Challenor Incident (March 2021)

Reddit hired Aimee Challenor (later Knight) as an admin. When an r/ukpolitics moderator posted an article mentioning her, they were permanently suspended. Reddit began auto-banning anyone mentioning her name. Over 200 subreddits went private in protest. Reddit CEO Huffman admitted she “wasn’t properly vetted” and she was fired.

DEFAMATION NOTE: All facts here are from Reddit’s own public statements and major news outlet coverage.


2. MaxwellHill

The Account

u/MaxwellHill was one of Reddit’s most prolific accounts — the first to surpass one million karma, ultimately accumulating nearly 15 million karma points. The account moderated influential subreddits including r/worldnews, r/travel, r/health, r/environment, and r/bad_cop_no_donut.

The Dormancy

The account’s last public post was July 1, 2020 — one day before Ghislaine Maxwell’s arrest on sex trafficking charges on July 2, 2020. The account has been publicly silent for over five years.

The Circumstantial Case

  • Name correspondence: “Maxwell” + “Hill” — Maxwell’s family estate in the UK was Headington Hill Hall

  • Posting gaps: Advocates claim posting gaps correlate with significant events in Ghislaine Maxwell’s life

  • Timing of dormancy: Account went permanently silent precisely at arrest

  • 2011 Gizmodo interaction: MaxwellHill mentioned being “busy with a potential business venture”

  • Source: Domi Good

  • Source: Film Daily

The Counter-Evidence

Reddit’s Response

Reddit has never confirmed or denied the identity of MaxwellHill. No official statement has been issued. The account remains dormant but not deleted. This silence is notable — Reddit has confirmed identities in other situations (celebrity AMAs, employee accounts).

DEFAMATION NOTE: The MaxwellHill theory remains unconfirmed. Both the circumstantial case and the counter-evidence are presented. Reddit’s non-response is documented fact. The book should present this as “documented timeline and Reddit’s silence” — not as established fact.


3. Documented Astroturfing Operations

Eglin Air Force Base

In 2013, Reddit’s own annual blog post revealed that Eglin Air Force Base was the “most addicted city (over 100k visits total).” The post was subsequently removed (archived). Eglin’s population was approximately 2,644 people — meaning each person would need to visit Reddit on roughly 38 separate accounts daily to account for the traffic. Eglin AFB is home to the 7th Special Forces Group and had published research papers on “Containment Control for a Social Network with State-Dependent Connectivity.”

DEFAMATION NOTE: The “most addicted city” statistic is from Reddit’s own blog. The military base population and PSYOP unit presence are public record. The research paper on social network containment is published. The connection between these facts is circumstantial.

Russian Internet Research Agency

Reddit banned 944 accounts linked to the Internet Research Agency in 2018 after Daily Beast reporting.

Iranian Operations

Community moderators identified Iranian propaganda operations on Reddit before Reddit’s admins did. Their warnings were initially ignored.

2024 U.S. Election Campaign

A report documented that the Harris-Walz campaign ran an organized astroturfing operation on Reddit: 126 of the top 1,000 posts in r/Politics over a one-month period were posted by official campaign volunteers. Volunteers made 2,551 posts garnering 5.7 million upvotes and 418,000 comments over 15 days.

DEFAMATION NOTE: The 2024 campaign astroturfing data comes primarily from right-leaning outlets (The Federalist, AllSides). The underlying data (post counts, karma, account identification) should be independently verified against the primary posts before publication. Reddit did not publicly address the allegations.

Scale

Over 300 manipulation operations identified on Reddit 2020–2024. Reddit’s transparency reports noted 1,200+ account suspensions annually for “inauthentic behavior.”


4. The Censorship Infrastructure Network

Stanford Internet Observatory (SIO)

  • Founded: 2019, Stanford’s Freeman Spogli Institute

  • Founding Director: Alex Stamos (former Facebook CISO)

  • Research Director: Renee DiResta

  • The Election Integrity Partnership (EIP): Co-led by DiResta, established July 2020. Submitted 4,800 flagged URLs to platforms with specific action recommendations including reducing “discoverability,” “suspending accounts’ ability to tweet,” and removing posts.

  • Murthy v. Missouri: EIP became central to the Supreme Court case. Preliminary injunction blocked government communication with EIP, Virality Project, and SIO.

  • Stamos departed November 2023; DiResta’s contract not renewed June 2024

  • Staff reduced to three employees by June 2024

  • Stanford’s statement: “SIO has not shut down or dismantled” but “faces funding challenges”

  • Cost to Stanford: “Two ongoing lawsuits and two congressional inquiries” cost “millions of dollars in legal fees”

  • House Judiciary Committee celebrated the dismantling as a “big win”

  • SIO announced it would not conduct research into the 2024 election or future elections

  • Source: Platformer

  • Source: Washington Post

  • Source: Stanford FSI

  • Source: Washington Times

  • Source: NYU Stern

  • Source: House Judiciary EIP Report

DEFAMATION NOTE: SIO’s own statement says it was not “dismantled.” The staff reduction to three is documented by Platformer and Washington Post. The House Judiciary Committee’s characterization as a “big win” and NYU’s characterization as a “dangerous partisan victory” are both cited. Present both framings.

Renee DiResta

  • BS in Computer Science and Political Science, Stony Brook University
  • CIA intern as an undergraduate (ended 2004 by her own account)
  • Led 2018 Senate investigation into Russian Internet Research Agency
  • Published Invisible Rulers: The People Who Turn Lies Into Reality (2024) (PublicAffairs)
  • Moved to Georgetown University’s McCourt School of Public Policy as associate research professor (October 2024)

The “CIA Renee” narrative: her undergraduate CIA internship was leveraged by critics (especially Mike Benz and Michael Shellenberger) to allege she was a CIA operative running censorship. Stamos described her in a 2019 livestream as having “worked for the CIA.” She calls this “baseless” and notes the internship ended before Twitter existed.

DEFAMATION NOTE: DiResta’s CIA internship is confirmed by her supervisor’s own public statement and her own acknowledgment. The characterization of her role at the CIA as “intern” vs. “operative” is contested. Both framings sourced. Her published work (Invisible Rulers) is publicly available for direct quotation.

Atlantic Council Digital Forensic Research Lab (DFRLab)

  • Founded: 2016, incubated at the Atlantic Council (NATO-aligned think tank)

  • Maintains the Foreign Interference Attribution Tracker (FIAT)

  • Funding: U.S. State Department, NATO, UAE, defense contractors (Raytheon, Lockheed Martin), tech companies

  • Source: DFRLab

Global Engagement Center (GEC)

  • Established: 2016 under Obama at the State Department

  • ~120 staff, $61 million annual budget

  • Defunded and closed December 23, 2024 after Congress declined to extend authorization

  • Funded or partnered with organizations accused of targeting domestic free speech (NewsGuard, GDI)

  • Elon Musk called the GEC the “worst offender in U.S. government censorship & media manipulation”

  • Source: CRS Report

  • Source: CyberScoop

  • Source: Fox News

Global Disinformation Index (GDI)

Compiles a “dynamic exclusion list” rating 2,000 websites for “disinformation” risk. Major advertising companies paid GDI for access and blacklisted high-risk entries.

Documented political bias: All 10 outlets GDI identified as “riskiest” leaned to the political right; all but one of the 10 ranked “least risky” leaned to the political left.

  • Received $100,000 from U.S. State Department via GEC

  • Microsoft-owned Xandr suspended its use of GDI pending internal review

  • Both British FCDO and U.S. State Department ceased funding GDI

  • Daily Wire and The Federalist filed lawsuits

  • Source: Wikipedia

  • Source: Washington Examiner

NewsGuard

  • Founded 2018 by Steven Brill and L. Gordon Crovitz

  • Provides “credibility ratings” for news websites

  • Received $25,000 grant from State Department’s Global Engagement Center

  • House Oversight probe: Chairman James Comer accused NewsGuard of anti-conservative bias

  • GEC shutdown (Dec 2024) effectively ended State Department support

  • Source: Wikipedia

  • Source: Capital Research Center

Center for Countering Digital Hate (CCDH)

  • Founded by Imran Ahmed (British-born, based in London and Washington D.C.)

  • Made “Kill Musk’s Twitter” an organizational priority for 2024

  • House Judiciary subpoena (August 2023): demanded documents on communications with federal agencies and platforms

  • X Corp. lawsuit dismissed (March 2024); X appealed to Ninth Circuit

  • December 2025: Trump administration attempted to deport Ahmed, accusing him of working to “censor freedom of speech.” A federal judge in Manhattan blocked the deportation. Ahmed is a lawful permanent resident with a U.S. citizen wife and son.

  • Source: JURIST

  • Source: The Hill

DEFAMATION NOTE: “Kill Musk’s Twitter” is documented from CCDH’s own organizational communications as reported by multiple outlets. The deportation attempt and judicial block are matters of public record.

ADL (Anti-Defamation League)

  • Annual “Online Hate and Harassment” report serves as industry benchmark

  • Social Media Scorecard: Rates platforms on hate speech response

  • Legislative advocacy: STOP HATE Act — would mandate social media companies work with the federal government on moderation policies

  • Shareholder activism: In 2025, ADL and JLens pressured Meta shareholders to vote for enhanced hate content reporting

  • Controversy: Criticized for conflating criticism of Israel with antisemitism. 2024 ADL-backed bill accused of threatening to “censor Israel criticism on social media”

  • CEO Jonathan Greenblatt accused social media rollbacks of causing an “explosion of hate”

  • Source: ADL 2024 Report

  • Source: STOP HATE Act

  • Source: Common Dreams

  • Source: Times of Israel

Trust & Safety Professional Association (TSPA)

  • Runs “All Things in Moderation” conference (2025 theme: “Third Spaces”)

  • Hosts the Trust & Safety Research Conference at Stanford

  • Publishes the Journal of Online Trust & Safety

  • Runs the Trust & Safety Teaching Consortium

  • Functions as the professional credentialing body for the T&S industry

  • Source: TSPA


5. The Revolving Door

PersonPathNotes
Renee DiRestaCIA intern → SIO Research Director → Georgetown“CIA Renee” controversy
Alex StamosFacebook CISO → SIO founding director → departed Nov 2023Left citing political pressure
Yoel RothTwitter Head of Trust & Safety → UC Berkeley / Carnegie EndowmentDeparted after Musk acquisition (Carnegie Endowment)
Mike BenzState Dept deputy assistant → Foundation for Freedom OnlineNBC reported alt-right content history
Imran AhmedCCDH founder → deportation target Dec 2025 → blocked by federal judge“Kill Musk’s Twitter” priority

DEFAMATION NOTE on Benz: NBC News reported Benz “appears to have been a pseudonymous alt-right content creator who courted and interacted with white nationalists and posted videos espousing racist conspiracy theories.” Benz responded this was part of “a deradicalization project” in which he participated in a “limited manner.” Both claims sourced. Source: https://www.nbcnews.com/tech/internet/michael-benz-rising-voice-conservative-criticism-online-censorship-rcna119213 (NBC News)


6. Government-Platform Coordination

Twitter Files (December 2022 – March 2023)

Internal Twitter documents released by journalists given access by Elon Musk.

Key findings:

  • Internal deliberations on Hunter Biden laptop story moderation; both Biden campaign and Trump White House flagged tweets for removal
  • “Visibility filtering” (shadow-banning): “Search Blacklist” for Dan Bongino, “Trends Blacklist” for Stanford’s Dr. Jay Bhattacharya, “Do Not Amplify” for Charlie Kirk
  • Twitter held regular industry meetings with FBI and DHS
  • A formal system existed for receiving “thousands of content reports from every corner of government including HHS, Treasury, NSA, and local police”

The Virality Project (Stanford-based): “routinely framed real testimonials about side effects as misinformation” including “true stories of blood clots from AstraZeneca vaccines.” Told Twitter that “true stories that could fuel hesitancy” should be considered “Standard Vaccine Misinformation.” Its final report claimed it was misinformation to suggest the vaccine does not prevent transmission or that governments were planning vaccine passports — both of which were true.

Disputed interpretation: Twitter attorneys denied the Files showed government coercion. Various journalists characterized the evidence as showing “little more than Twitter’s policy team struggling with difficult decisions.”

DEFAMATION NOTE: Present the documented facts (meetings occurred, spreadsheets sent, Virality Project flagged true stories) alongside the disputed interpretation of what those facts mean. The facts themselves are not contested. The characterization of those facts as “coercion” vs. “coordination” vs. “difficult decisions” is.

CISA Switchboarding

CISA conducted “switchboarding” — flagging accounts or posts on social media to platforms for moderation review. CISA confirmed it “did not provide the switchboarding service for the 2022 election cycle and has no intention to engage in switchboarding for the next election.”

Murthy v. Missouri (Supreme Court, June 26, 2024)

Missouri, Louisiana, and five social media users alleged that government communications urging platforms to act against COVID-19 and 2020 election “misinformation” violated the First Amendment.

Ruling: 6-3, Justice Barrett writing. Plaintiffs lacked standing — none could show government statements were the likely cause of specific platform actions against them. The Court avoided ruling on the merits.

Alito’s dissent (joined by Thomas and Gorsuch): “This is one of the most important free speech cases to reach this Court in years.”

Moody v. NetChoice (Supreme Court, July 1, 2024)

Unanimous 9-0 (Justice Kagan). Confirmed that social media platforms are protected by the First Amendment when exercising editorial discretion to select, organize, display, promote, demote, or block content — even through algorithms.

EU Digital Services Act (DSA)

Fully applicable since early 2024 for all “Very Large Online Platforms” (45M+ monthly EU users). In H1 2025, platforms reported more than 9 billion content moderation decisions, with 99% taken proactively (not in response to user reports). Out-of-court settlement bodies overturned platforms’ decisions in 52% of closed cases. In two years, platforms reversed almost 50 million decisions.

  • Source: EU

7. AI Moderation Replacing Human Gatekeepers

Platform Transitions

  • Meta: Removed 26 million pieces of hate speech in one quarter, 97% detected by AI before any user flagged them. In February 2024, received 7 million appeals from content removals — 8 out of 10 appealed, with 1 in 3 saying “it was a joke.” Meta acknowledged 1-2 out of every 10 removal actions may have been mistakes.

  • TikTok: Laid off 700 human moderators in 2024 in favor of AI systems (Vice)

  • YouTube: Creators report daily wrongful channel terminations by automated systems. AI has banned original creators while leaving up stolen re-uploads.

  • X/Twitter: ~73% of removed tweets first flagged by AI. Trust & safety staff cut from 279 engineers to 55 (80% reduction). Content moderation team cut from 107 to 51.

Meta’s January 2025 Reversal

Mark Zuckerberg announced Meta would:

  • End third-party fact-checking program in the US, replacing with Community Notes

  • Scale back automated moderation to focus on “high severity violations” only (terrorism, CSAM, drugs, fraud)

  • Move Trust & Safety teams from California to Texas

  • Relax content restrictions on “immigration, gender and other hot-button issues”

  • Source: Meta

  • Source: NPR

EFF response: Rather than addressing historically over-moderated subjects, Meta made “targeted changes to its hateful conduct policy that would allow dehumanizing statements to be made about certain vulnerable groups.”

AI Content Moderation Market

Projected to grow from $1.03 billion (2024) to $2.59 billion (2029). Companies like Checkstep claim 90% automation rates.

The T&S Collapse (2023–2025)

  • X: Cut trust & safety roles by 43%, safety engineers by 80%, global public policy by 80%
  • Meta: 21,000-job layoffs had “outsized effect” on T&S work
  • Google: Cut ~33% of its misinformation/radicalization unit
  • CHI 2025 paper: “The End of Trust and Safety?”

8. Dead Internet / Bot Traffic

  • Imperva 2024-2025: Automated bot traffic surpassed human-generated traffic for the first time, constituting 51% of all web traffic in 2024. Bad bots alone: 37% of all internet traffic.

  • Originality.AI (2025): 15% of Reddit posts are likely AI-generated (up from 13% in 2024). From 2021 to 2024, AI-generated content increased by 146.30%. In r/NoSleep: 41.39% AI content.

  • Sam Altman (September 3, 2025): “i never took the dead internet theory that seriously but it seems like there are really a lot of LLM-run twitter accounts now”

  • Facebook “AI slop” (2024): AI-generated images including “Shrimp Jesus,” fake children’s artwork, and flight attendant images went viral and received millions of engagements. (The Conversation)

  • Academic survey (2025): “The Dead Internet Theory: A Survey on Artificial Interactions and the Future of Social Media”


9. Counter-Voices

FIRE (Foundation for Individual Rights and Expression)

Nonpartisan 501(c)(3), founded 1999.

Three core principles:

  1. Law should require transparency whenever government involves itself in social media moderation
  2. Content moderation policies should be transparent to users with appeals
  3. Moderation decisions should be unbiased and consistently applied

“Americans don’t trust the government to make social media content decisions.” Even when platforms’ decisions are “unwise or biased,” empowering government to dictate moderation would be “a cure far worse than the disease.”

EFF (Electronic Frontier Foundation)

Nuanced position — opposes both government coercion AND platform over-moderation:

  • “Government coercion resulting in censoring users’ speech online violates the First Amendment”

  • Social media platforms have First Amendment rights to curate third-party speech

  • Co-authored the Santa Clara Principles (transparency guidelines for moderation)

  • “Censorship broadly is not the answer to misinformation”

  • Source: EFF

Nadine Strossen (former ACLU president)

“No matter how great the potential harm of the speech, the potential harm of censorship is even greater.”

Eugene Volokh (Hoover Institution, Stanford)

Published extensively on treating social media platforms as common carriers. 2025 joint statement with constitutional scholars: “the government may not threaten funding cuts as a tool to pressure recipients into suppressing First Amendment-protected speech.”


10. The Collapse Timeline

DateEvent
Nov 2023Stamos departs SIO
Jun 2024SIO reduced to 3 staff, DiResta not renewed
Jun 2024Murthy v. Missouri dismissed on standing (6-3)
Jul 2024Moody v. NetChoice: platforms have 1A editorial discretion (9-0)
2023-2024X cuts T&S by 80%; Meta, Google make significant cuts
Dec 2024Global Engagement Center defunded and closed
Dec 2024GDI defunded by US and UK governments
Jan 2025Meta ends fact-checking, scales back automated moderation
Dec 2025Reddit implements moderator cap (triggered by r/Art meltdown)
Dec 2025CCDH founder Ahmed faces deportation, blocked by federal judge

The infrastructure is being dismantled. The people are relocating, not retiring. DiResta moved to Georgetown. Stamos remains at Stanford. Roth is at Berkeley and Carnegie Endowment. The question for Book 3: what happens when the same instinct — control discourse through institutional gatekeeping — migrates from human moderation networks to AI training data and content curation algorithms? That migration is the through-line of the Evil Robots series, which tracks the machine successor in AI-Powered Content Moderation.


11. The Architectural Kill Shot: Prompt Injection and the End of AI Moderation

The Impossibility Results

All three frontier labs and a national intelligence agency have acknowledged that prompt injection — the ability to override AI safety instructions through crafted inputs — is likely a permanent, unfixable vulnerability.

OpenAI (December 22, 2025) “Prompt injection, much like scams and social engineering on the web, is unlikely to ever be fully ‘solved.’” Published in a security update for ChatGPT Atlas after automated red-teaming found injection attacks that could make an AI browser agent send a resignation letter when asked to draft an out-of-office reply.

UK National Cyber Security Centre (December 8, 2025) NCSC (part of GCHQ) published “Prompt Injection Is Not SQL Injection” arguing the vulnerability is structurally worse than SQL injection. SQL injection was fixable because databases separate data from instructions. LLMs cannot: “Under the hood of an LLM, there’s no distinction made between data or instructions; there is only ever ’next token.’” Conclusion: prompt injection “may never be totally mitigated.” If security is critical, “it may not be a good use case for LLMs.”

Anthropic (November 24, 2025) “Prompt injection is far from a solved problem, particularly as models take more real-world actions.” Their best result — 1% attack success rate — described as “meaningful risk rather than a solved problem.” Claude Cowork was exploited via prompt injection within 48 hours of launch (January 2026). The vulnerability had been disclosed via HackerOne three months prior.

The Empirical Confirmation

“The Attacker Moves Second” (October 2025) Joint paper by researchers from Google DeepMind, OpenAI, and Anthropic. Tested 12 published AI defenses that originally claimed near-zero attack success. Using adaptive attacks (gradient descent, RL, random search, human-guided), all 12 defenses bypassed above 90%. Prompting-based defenses: 95-99% bypass. Training-based defenses: 96-100% bypass.

  • Source: Nasr, Carlini, Sitawarin, Schulhoff et al. arXiv:2510.09023 (October 2025)
  • Coverage: VentureBeat
  • Coverage: Simon Willison

Cisco “Death by a Thousand Prompts” (December 2025) AI models block 87% of single-shot attacks. Under conversational persistence (probing, reframing, escalating), block rate collapses to 8%. Attack success: 13% → 92% with persistence.

  • Source: Chang, Conley, Ganesan, Swanda. Cisco AI Threat Research (December 1, 2025)
  • Coverage: VentureBeat

GPT-5 Jailbroken in Under 24 Hours (August 2025) Released August 7, 2025 with new “safe-completion” method. Three independent teams (NeuralTrust, SPLX, Tenable) bypassed it within 24 hours. NeuralTrust: “Echo Chamber” jailbreak, Molotov cocktail instructions after 3 prompts. Tenable: Crescendo technique, 4 prompts.

OWASP LLM01:2025 — Prompt Injection as #1 Vulnerability “Prompt injection vulnerabilities are possible due to the nature of generative AI. Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention.”

Community Notes: The Replacement That Depends on What It Replaced

ACL 2025 (University of Copenhagen) Community Notes cite fact-checking sources up to 5x more than previously reported. Posts linked to misinformation narratives are 2x as likely to reference professional fact-checkers. Conclusion: “successful community moderation relies on professional fact-checking” — the systems are “deeply intertwined,” not interchangeable.

Community Notes Effectiveness (PNAS 2025) Reposts dropped 46%, likes dropped 44% after notes applied. Posts with notes 32% more likely to be deleted by authors. Medical professionals rated 98% of COVID notes accurate.

Critical Limitations Only 11% of submitted notes reach “helpful” status and are shown to users. Average display time: 24.29 hours (vs. debunking comments appearing within 2 hours). X’s system has been infiltrated by ideological blocs coordinating up/downvotes.

Simon Willison (coined “prompt injection,” September 2022)

Core insight: the vulnerability is architectural. No privilege separation in LLMs, no data/control path separation. Delimiters can be forged, instruction hierarchies can be circumvented, separate models double the attack surface. Every proposed solution introduces new injection vectors.

Bluesky’s Stackable Moderation (Alternative Architecture)

Bluesky grew from 25.94M to 41.41M users in 2025. Their approach: human-centered moderation for context-dependent decisions, AI only for spam/CSAM/coordinated attacks. Open-sourced Ozone labeling service. “Stackable moderation” — users layer multiple moderation services. Protocol-level design separates moderation from hosting, identity, and algorithmic feed.


Source URLs