This study examines the challenges involved in automatically identifying disinformation about abortion on Twitter. We collected 166,180 tweets posted on Twitter (January–December 2022) about the Supreme Court’s reversal of Roe v. Wade to train large language models (LLMs) and machine learning systems to recognize disinformation about abortion. For this purpose, we created a pilot corpus of 8,309 tweets. Surprisingly, only 0.08% of the pilot corpus contained medical disinformation. We thus questioned whether the planned machine learning could be carried out, given that semantic ambiguity regarding what is fake in the abortion debate poses significant hurdles. We observed that tweets expressing extreme viewpoints were often labelled as “fake”, highlighting the subjective nature of such categorizations. These findings, strongly influenced by personal ideological views on the topic in question, emphasize the complexity of using LLMs and machine learning systems to navigate emotionally charged topics. They also stress on the importance of considering different perspectives to reduce bias in analyses, advise caution against relying solely on these technologies, and warn of the problems of potential bias and “cherry-picking” in data interpretation, especially when researching social media debates which are full of implicit content, presuppositions, and implicatures, as well as fallacious argumentation.

The researcher’s bias in fake news automatic detection: a case study

Maci, Stefania Maria
;
2026

Abstract

This study examines the challenges involved in automatically identifying disinformation about abortion on Twitter. We collected 166,180 tweets posted on Twitter (January–December 2022) about the Supreme Court’s reversal of Roe v. Wade to train large language models (LLMs) and machine learning systems to recognize disinformation about abortion. For this purpose, we created a pilot corpus of 8,309 tweets. Surprisingly, only 0.08% of the pilot corpus contained medical disinformation. We thus questioned whether the planned machine learning could be carried out, given that semantic ambiguity regarding what is fake in the abortion debate poses significant hurdles. We observed that tweets expressing extreme viewpoints were often labelled as “fake”, highlighting the subjective nature of such categorizations. These findings, strongly influenced by personal ideological views on the topic in question, emphasize the complexity of using LLMs and machine learning systems to navigate emotionally charged topics. They also stress on the importance of considering different perspectives to reduce bias in analyses, advise caution against relying solely on these technologies, and warn of the problems of potential bias and “cherry-picking” in data interpretation, especially when researching social media debates which are full of implicit content, presuppositions, and implicatures, as well as fallacious argumentation.
2026
2026
File in questo prodotto:
File Dimensione Formato  
10.1515_lingvan-2024-0060.pdf

non disponibili

Descrizione: Articolo su rivista
Tipologia: Versione dell'editore
Licenza: Accesso chiuso-personale
Dimensione 556.34 kB
Formato Adobe PDF
556.34 kB Adobe PDF   Visualizza/Apri

I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10278/5127391
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
social impact