Literature scan & annotation

Explore the data

Analyses beyond the main browse table, over the combined keyword-search + forward/backward-snowball corpus (22,165 records, 3,233 included). ← back to the main review

Ask the data

Answers computed live from the same data as the charts below -- institutions, countries, venues, bot personas, governance categories, years, and citation counts. No per-author data yet (only institution/country affiliations were fetched), so a question about a specific person won't resolve, only the institution named in it. Try: "which institution has the most papers", "how many papers from Carnegie Mellon University", "most cited paper", "papers about chatbots in 2023", "which country has the most citations".

Seed-centered citation network

The 7 seed papers (large diamonds), each shown with its top 18 highest-cited backward references and forward citing papers (colored by which seed they came from; a paper connected to two seeds shows only its first-found link here). Solid seed-to-seed lines mean one seed cites another directly. Hover a node for its title; node size scales with citation count.

Observation: 28 direct citation links exist among the 7 seeds themselves, and the single highest-cited neighbor pulled in by any seed is “The spread of true and false news online” (8,836 citations).

Seed-pair bibliographic coupling & co-citation

Bibliographic coupling: how many references two seeds share (a proxy for topical closeness). Co-citation: how many later papers cite both seeds together. Darker = more overlap.

Observation: The two seeds with the most in common are “Botometer 101: social bot practicum for computatio” and “A global comparison of social media bot and human ”, sharing 12 references (bibliographic coupling) — more than any other seed pair.

Shared references (bibliographic coupling)
Shared citing papers (co-citation)

Who publishes this: institutions & countries

Author affiliations for included papers only (excluded papers aren't attributed to institutions). "Publications" counts a paper once per institution/country that appears on it (a paper with two US co-authors from different schools counts once for the US, once for each school). "Citations" sums each paper's OpenAlex cited-by count into every institution/country on it -- a proxy for whose work in this space gets picked up, not a per-institution citation count in the strict bibliometric sense.

Observation: United States leads by publication count (784 papers), while United States leads by total citations (51,342), and the single most-published institution is Carnegie Mellon University.

Top institutions
Top countries

Where this literature is published

Top 20 venues by record count, across the full combined corpus (included + excluded).

Observation: arXiv (Cornell University) is the single most common venue, with 679 records — among 20 distinct venues shown here alone.

Citation counts

Distribution of OpenAlex cited-by counts across all 22,165 records, and the 20 most-cited papers in the corpus.

Observation: The single most-cited RECORD overall is “Deep Residual Learning for Image Recognition” (229,732 citations) — excluded at screening (not actually about social media bots); among INCLUDED papers, the most-cited is “Comparing Physician and Artificial Intelligence Chatbot Responses to P…” (2,611 citations, 2023).

Top 20 most-cited papers

Citation velocity

Citations per year since publication (cited-by count ÷ years since publication), not raw cited-by count -- so a 2018 paper with 200 citations and a 2024 paper with 50 aren't compared unfairly. Each dot is one included paper; hover for its title. The line is the average velocity per publication year.

Observation: Average citation velocity has fallen from 8.2/yr in 2016 to 0.21/yr in 2026; the fastest-accumulating paper is “Comparing Physician and Artificial Intelligence Chatbot Responses to P…” at 652.75 citations/year.

Top 20 by citation velocity

Core citation graph

Edges BETWEEN included papers only -- who in this literature cites whom -- as opposed to the seed-centered network above (which only shows the 7 seeds' immediate neighbors). Shows the 150 most internally-cited papers (out of 14992 internal citation edges total among all included papers); node size and ring position scale with in-degree (how many other included papers cite it). Hover for title.

Observation: “The rise of social bots” is cited by 493 other included papers — the most of any paper in this corpus — out of 14,992 internal citation edges total.

Most internally-cited papers

Topic drift (title keywords, by year)

Top keyword/phrase frequency extracted from included papers' TITLES only (abstracts aren't kept in the published dataset), by publication year -- a local, fully computed-from-title-text proxy for topic drift, not OpenAlex's own topic classification. Common English stopwords and corpus-constant terms ("social", "media", "bot(s)", "detection", etc.) are filtered out so the terms that DO show up are the ones that distinguish one year from another.

Observation: “learning” shows the largest rise of any tracked title-keyword, from 0 mentions in 2016 to 37 in 2026.

Bot-persona co-occurrence

Among included papers, how often each pair of the 10 most-common bot personas is tagged on the SAME paper (diagonal = total papers with that persona).

Observation: Unclassified is the most-tagged bot persona overall (1460 papers), and Social Influence Bot + Amplifier Bot co-occur on the same paper more than any other pair (72 papers).

Governance actor × typology

Among included papers, which governance actor (platform / government / civil society) each governance-typology idea is addressed to.

Observation: The most common governance pairing is government actors addressed via Legal & Regulatory Frameworks (156 papers).

Search vs. snowball, over time

How much of each publication year's records came from the keyword-search stream vs. the forward/backward citation snowball. Snowball papers skew slightly older on average, since they're reachable by citation from mostly-older seed papers.

Observation: The forward/backward snowball's share of each year's records has shrunk from 92% in 2016 to 78% in 2026.