These results not only highlight but also provide clear evidence that the fundamental risk of manipulation or inflation of metrics by Twitter bots, whether deliberate or not, within altmetrics is a genuine concern
Español 🇪🇸
Uno de los múltiples reproches hacia las altmétricas ha sido su vulnerabilidad ante la posible manipulación, especialmente mediante la presencia de bots en Twitter (ahora X). Este interrogante se ha mantenido desde sus inicios y, con la llegada de Elon Musk y los sucesivos cambios en el acceso a la plataforma, el escenario se volvió aún más complejo.
Fue en septiembre de 2022 cuando llegó a mis manos el artículo Botometer 101: social bot practicum for computational social scientists, el cual revisa en detalle el uso de Botometer para la investigación. Rápidamente se lo comenté a Dan y se nos ocurrió aplicar ese enfoque para arrojar un poco más de luz a tan clásica cuestión: ¿quiénes son realmente los usuarios que mencionan papers científicos en Twitter? Si bien en el trabajo se mostraban las utilidades y limitaciones de la herramienta, observamos que, en nuestro ámbito, a diferencia de entornos más complejos como la esfera política, donde los bots intentan parecer humanos, abundan cuentas automatizadas que funcionan como alertas o notificaciones, sin intención de ocultar su naturaleza. Aunque su propósito no sea malicioso, su presencia masiva puede distorsionar los indicadores altmétricos y ofrecer una imagen irreal del impacto social de ciertas publicaciones.

Durante casi tres meses nos dedicamos intensamente a estudiar cómo construir el marco metodológico y acceder a los datos de Botometer mediante su API. Este trabajo fue posible finalmente gracias al respaldo del proyecto InfluScience, desde donde ahora escribimos, que nos permitió financiar los casi 700 dólares necesarios para consultar los perfiles de las más de 4,8 millones de cuentas de Twitter que conformaban nuestro dataset. Se trataba de usuarios que habían mencionado cerca de 3,7 millones de artículos científicos publicados entre 2017 y 2021, generando más de 51 millones de menciones.
Exploramos diversas estrategias para identificar cuentas automatizadas, combinando enfoques cuantitativos con criterios más cualitativos. Finalmente, diseñamos un método híbrido que resultó ser altamente eficiente. Este se basaba en el uso de BotometerLite, una versión simplificada pero eficaz del Botometer original (de haber usado la versión completa habríamos necesitado más tiempo y presupuesto, y el resultado final no habría diferido sustancialmente), combinado con parámetros de actividad de las cuentas de Twitter, como el número total de menciones o la frecuencia de publicación. Para afinar la detección, realizamos además una validación manual exhaustiva sobre una muestra estratificada de 1.180 cuentas, lo que nos permitió ajustar los umbrales y minimizar los falsos positivos. El modelo final identificó como bots aquellas cuentas con un botscore superior a 0,6 y más de 8 menciones, alcanzando una precisión muy alta (F1 ≈ 0,83). La validación específica en disciplinas como Matemáticas confirmó la robustez del modelo: más del 80% de las cuentas clasificadas como bots realmente lo eran.
Con esta base, elaboramos un primer análisis exploratorio que presentamos en el congreso STI 2023, centrado especialmente en el área de Social Sciences y con atención detallada a Information Science & Library Science. En aquel momento optamos por un título provocador y algo irónico: The Elon Musk Paradox: Quantifying the Presence and Impact of Twitter Bots on Altmetrics with Focus in Social Sciences.
A pesar del entusiasmo inicial, el estudio quedó parcialmente aparcado durante los meses siguientes por la recta final de mi tesis doctoral. No obstante, logramos retomarlo a finales de 2023, preparar el manuscrito completo y publicar el preprint. Primero lo enviamos a Information Processing & Management, pero fue rechazado sin pasar a revisión. Posteriormente lo remitimos a JASIST, donde, tras un largo proceso de evaluación que se prolongó casi un año, finalmente fue aceptado y publicado.

Ahora os presentamos la versión final del estudio, que ofrece una respuesta matizada a esa vieja pregunta: ¿los bots distorsionan las altmétricas? Y, como ocurre tantas veces en ciencia, la respuesta es: depende. Si bien los bots representan solo el 0,23% de los usuarios, son responsables de un 11,5% de los tuits totales sobre artículos científicos. Su impacto es mínimo en áreas como Ciencias Sociales o Humanidades, donde apenas suponen un 4% de los tuits, pero en disciplinas como Matemáticas, su presencia es mucho más relevante. En promedio, el 44,5% de los tuits en esta área provienen de bots, y si descendemos al nivel de especialidad, encontramos subcategorías como Mathematics o Applied Mathematics donde los bots generan hasta el 70% de los tuits, e incluso son responsables del 100% de la visibilidad social de muchos artículos. Al final, cada área tiene su propia lógica de comunicación científica, y aquellas con menor presencia humana en Twitter—pero alta visibilidad en repositorios como arXiv—son las más afectadas. Un recordatorio de que, incluso en altmetrics, el contexto lo es todo.
REFERENCIA DEL PAPER
Arroyo-Machado, W., Herrera-Viedma, E., & Torres-Salinas, D. (2025). The botization of science? Large-scale study of the presence and impact of Twitter bots in science dissemination. Journal of the Association for Information Science and Technology, 1–18. https://doi.org/10.1002/asi.24998
English 🇺🇸
One of the most recurring criticisms of altmetrics has been their vulnerability to manipulation, particularly due to the presence of bots on Twitter (now X). This concern has been present since the early days of social media metrics, and with Elon Musk’s acquisition of the platform and the changes that followed, the scenario became even more opaque.
In September 2022, I came across the paper Botometer 101: social bot practicum for computational social scientists, which provided a detailed review of how to use Botometer in research. I immediately shared it with Dan, and we saw an opportunity to revisit an old but unresolved question: who are the users that mention scholarly papers on Twitter? While the paper showed both the potential and the limitations of Botometer, we knew that our case was different. Unlike the political domain —where bots often mimic human behavior— the academic landscape includes numerous automated accounts that openly act as alert systems, without trying to conceal their nature. And while their intent is not malicious, their massive presence can seriously distort altmetric indicators and create a misleading picture of societal impact.

For nearly three months, we worked intensively on designing the methodological framework and gaining access to Botometer’s data via its API. This endeavor was made possible thanks to the InfluScience project, which supported the nearly $700 needed to analyze over 4.8 million Twitter accounts. These users had mentioned around 3.7 million scholarly papers published between 2017 and 2021, generating more than 51 million Twitter mentions.
We explored different strategies to identify automated accounts, combining quantitative thresholds with qualitative assessments. Eventually, we designed a hybrid method that proved highly efficient. It relied on BotometerLite, a simplified yet effective version of the original Botometer (using the full version would have required more time and funding, with minimal added benefit in our context), together with indicators of Twitter activity such as total mentions and tweet frequency. To fine-tune detection, we carried out a manual validation on a stratified sample of 1,180 accounts, which helped us calibrate the thresholds and reduce false positives. The final model classified as bots those accounts with a botscore above 0.6 and more than 8 mentions, achieving very high precision (F1 ≈ 0.83). Validation in fields like Mathematics confirmed its robustness: over 80% of accounts flagged as bots were correctly identified.
Based on this, we carried out a first exploratory study, which we presented at STI 2023, focusing on Social Sciences and in particular Information Science & Library Science. At the time, we chose a provocative and somewhat ironic title: The Elon Musk Paradox: Quantifying the Presence and Impact of Twitter Bots on Altmetrics with Focus in Social Sciences.
Although the initial momentum was strong, the study had to be temporarily set aside due to the final phase of my PhD dissertation. Eventually, we returned to it at the end of 2023, prepared the full manuscript, and released the preprint. It was first submitted to Information Processing & Management, where it was desk-rejected, and then to JASIST, where —after nearly a year of review— it was finally accepted and published.

We are now pleased to share the final version of this study, which offers a nuanced answer to an enduring question: do bots distort altmetrics? As is often the case in science, the answer is: it depends. While bots represent just 0.23% of all users, they are responsible for 11.5% of all tweets about scientific publications. Their impact is minimal in fields like Social Sciences or Humanities, where they account for only 4% of tweets, but in disciplines such as Mathematics, their influence is significantly higher. On average, 44.5% of tweets in this field are generated by bots, and if we zoom into specific specialties like Mathematics or Applied Mathematics, bots account for up to 70% of the tweets, and in many cases are responsible for 100% of a paper’s social visibility. Ultimately, each discipline has its own communication dynamics, and those with lower human presence on Twitter—but high reliance on repositories like arXiv—are the most affected. A reminder that, even in altmetrics, context is everything.
PAPER’S REFERENCE
Arroyo-Machado, W., Herrera-Viedma, E., & Torres-Salinas, D. (2025). The botization of science? Large-scale study of the presence and impact of Twitter bots in science dissemination. Journal of the Association for Information Science and Technology, 1–18. https://doi.org/10.1002/asi.24998