Data Science & Digitalization
Chair for Data Science & Digitalization
- News
-
BWL XI: Paper in Nature CommunicationsA new article has been accepted for publication in Nature Communications (IF: 15.7). In this work, we perform a large-scale quasi-experimental study to analyze whether community fact-checks reduce the spread of misleading posts on the social media platform X (formerly Twitter).
-
BWL XI: Paper in EPJ Data ScienceA new research paper studying political communication on TikTok has been accepted for publication in EPJ Data Science.
-
BWL XI: Two papers accepted at WWWTwo new research papers have been accepted for publication in the Proceedings of the ACM Web Conference (WWW '26). WWW is a premier publication outlet in data science with a low acceptance rate (CORE Ranking A*).
-
BWL XI: Paper in Nature's Scientific ReportsA new research paper has been accepted for publication in Nature's Scientific Reports. In this study, we empirically investigate the helpfulness of the context provided in community-created fact-checks on the social media platform X (formerly Twitter).
-
BWL XI: Study on Deepfakes at IC2S2Our study "Characterizing Deepfakes on X" has been accepted for presentation at the International Conference on Computational Social Science (IC2S2 '25).
-
BWL XI: Media Coverage in TIME Magazine & The AtlanticOur research on community-based fact-checking has been featured in TIME Magazine and The Atlantic.
-
BWL XI: Paper accepted in PNAS NexusA new research paper has been accepted for publication in PNAS Nexus. In our study, we estimate the link between online political advertising and election outcomes during the 2021 German federal election.
-
BWL XI: DFG Grant for Research on Community-Based Fact-CheckingThe German Research Foundation (DFG) has awarded a new research grant to Prof. Dr. Nicolas Pröllochs. The funding will support our research on community-based fact-checking on social media.
-
BWL XI: Paper in Nature Reviews PsychologyA new article has been accepted for publication in Nature Reviews Psychology (IF: 16.8). Together with an interdisciplinary team of domain experts, we describe how natural language processing (NLP) can be used to analyse text data in behavioural science.
- Contact Person
-
- Featured Research: Community-Based Fact-Checking Reduces the Spread of Misleading Posts on Social Media
-

Social media platforms increasingly rely on community-based fact-checking systems such as X’s Community Notes to combat misinformation at scale. In this study, we analyze more than 431 million reposts across 237,180 fact-checked cascades and provide large-scale causal evidence that community notes reduce the subsequent spread of misleading posts by 61.2% on average. We further find that community notes increase the likelihood that users delete misleading posts by 94.3%. However, notes often appear too late to prevent the early, most viral stage of diffusion, limiting their overall system-wide impact. Our findings highlight both the promise and current limitations of community-based fact-checking systems in reducing misinformation on social media.
- Research paper at Nature Communications (open access)
- Interview with FAZ
- Media Coverage (selection): The Washington Post, The Atlantic, TIME Magazine, ABC News, Poynter, BBC
- Featured Research: Negativity Drives Online News Consumption
-

Online media is important for society in informing and shaping opinions, hence raising the question of what drives online news consumption. Here, we analyze the causal effect of negative and emotional words on news consumption using a large online dataset of viral news stories. Specifically, we conducted our analyses using a series of randomized controlled trials (N = 22,743). Our dataset comprises ∼105,000 different variations of news stories from Upworthy.com that generated ∼5.7 million clicks across more than 370 million overall impressions. Although positive words were slightly more prevalent than negative words, we found that negative words in news headlines increased consumption rates (and positive words decreased consumption rates). For a headline of average length, each additional negative word increased the click-through rate by 2.3% Our results contribute to a better understanding of why users engage with online media.
- Research paper at Nature Human Behaviour
- Media coverage (selection): ARD, FAZ, Heise, Deutschlandfunk, ORF, Psychology Today, The Atlantic
News
A new research paper has been accepted for publication in the proceedings of 17th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2019). NAACL-HLT is one of the ACL flagship conferences in natural language processing (acceptance rate of ~20%).
A new research paper on fake news prevention has been accepted for presentation and inclusion in the conference proceedings at NeuroIS Retreat 2019, Vienna, Austria.
A new research grant from Google supports our research by funding cloud resources for machine learning applications in electronic commerce and finance.
A new software paper has been accepted for publication in Journal of Open Source Software. The paper presents the first R package for performing model-free reinforcement learning in R.
A new education grant from Google supports our teaching activities in the next semester by funding cloud resources for the course "Text Mining". The Google Education Grant provides each course participant with education credits that allow accessing cloud resources and data science tools from the Google Cloud Platform.
We will offer a master's course on "Text Mining" in winter semester 19/20. The course will have an interactive format including coding sessions, discussions, and presentations by students. The course is also open to interested bachelor students currently enrolled in the 210- and 240-CP programs.
A new research paper has been accepted for publication in the proceedings of the 40th International Conference on Information Systems (ICIS 2019). ICIS is the flagship conference in Information Systems and ranked A in the VHB ranking.
In summer semester 2020, we offer the course "Data Science for Management" for bachelor's students. Master's students are welcome to apply for a spot in the "Data Science Seminar." Seminar participants are welcome to propose seminar topics based on their personal interests. The first deadline for seminar applications is 27th January 2020 (via e-mail).
A new research paper has been accepted for publication in Expert Systems with Applications (IF: 4.292). The paper uses machine learning and methods from natural language processing to predict sentence-level polarity labels in financial news.
A new research paper on fake news has been accepted for presentation and inclusion in the conference proceedings at NeuroIS Retreat 2020, Vienna, Austria.
A new research paper on negation scope detection has been accepted for publication in Information Sciences (IF: 5.524).
A new workshop paper has been accepted for presentation at the SIGBPS Workshop on Blockchain and Financial Analytics 2020 at AMCIS 2020. The paper makes use of state-of-the-art methods in visual analytics and computer vision to learn price-relevant aesthetics of floor plans on online real estate platforms.
In winter semester 20/21, we offer the course "Text Mining" for master's students. Bachelor's students are welcome to apply for a spot in the "Data Science Proseminar." Proseminar participants are welcome to propose seminar topics based on their personal interests.
A new workshop paper has been accepted for presentation at the Conference on Digital Experimentation at MIT (CODE@MIT). The paper analyzes the effect of emotions in online clickbait on engagement rates.
A new workshop paper has been accepted for presentation at the WITS 2020. The paper uses deep learning to integrate price-relevant aesthetics of floor plans into hedonic price models.
Prof. Dr. Nicolas Pröllochs has received a research grant from the German Research Foundation (DFG). The research grant supports our research on Covid-19-related misinformation on social media.
A new research paper has been accepted for publication in the proceedings of The Web Conference (WWW). The Web Conference is one of the flagship conferences in data science with a very low acceptance rate (CORE Ranking A*).
Kirill Solovev presented his research on machine learning for real estate markets at the Machine Learning Summer School (MLSS) 2021 in Taipei (virtual event).
Research
Our research focuses on the application of computational techniques for understanding and predicting human behavior on digital platforms. Current research projects leverage data science methods, and machine learning to drive domain-specific decisions in a broad selection of business-relevant areas, including, but not limited to, data analytics for social media and electronic commerce, financial data science, and natural language processing for business applications.
Teaching
Our focus in academic teaching is on courses at the interface between management science and computer science. Lectures and exercises are designed to provide students with strong quantitative skills that form the basis for a profound understanding of data science methods and data-driven decision making. Bachelor’s and master’s theses are typically embedded in our own research context and pave the way for students' own research efforts. Hands-on supervision enables students to already achieve meaningful successes and find enthusiasm for research.
Our focus in academic teaching is on courses at the interface between management science and computer science. Lectures and exercises are designed to provide students with strong quantitative skills that form the basis for a profound understanding of data science methods and data-driven decision making. Bachelor’s and master’s theses are typically embedded in our own research context and pave the way for students' own research efforts. Hands-on supervision enables students to already achieve meaningful successes and find enthusiasm for research.
Featured Research: Argumentation Patterns in Reviews
An overwhelming majority of previous works find longer reviews to be more helpful than short reviews. In this study, we propose that longer reviews should not be assumed to be uniformly more helpful; instead, we argue that the effect depends on the line of argumentation. To test this idea, we use a large dataset of customer reviews from Amazon in combination with a state-of-the-art approach from natural language processing that allows us to study argumentation lines at sentence level. Our results disprove the prevailing narrative that longer reviews are uniformly perceived as more helpful and allow retailer platforms to feature more useful product reviews.
Preprint on arXiv
Featured Research: Community-Based Fact-Checking on Twitter’s Birdwatch Platform
-

Twitter has recently introduced “Birdwatch,” a community-driven approach to address misinformation on Twitter. In this work, we empirically analyze how users interact with this new feature. Our empirical analysis yields the following main findings: (i) users more frequently file Birdwatch notes for misleading than not misleading tweets. These misleading tweets are primarily reported because of factual errors, lack of important context, or because they contain unverified claims. (ii) Birdwatch notes are more helpful to other users if they link to trustworthy sources and if they embed a more positive sentiment. (iii) The helpfulness of Birdwatch notes depends on the social influence of the author of the fact-checked tweet. For influential users with many followers, Birdwatch notes yield a lower level of consensus among users and community-created fact checks are more likely to be seen as being incorrect. Altogether, our findings can help social media platforms to formulate guidelines for users on how to write more helpful fact checks. At the same time, our analysis suggests that community-based fact-checking faces challenges regarding biased views and polarization among the user base.
Paper at ICWSM
Featured Research: Hate Speech in the Political Discourse on Social Media: Disparities Across Parties, Gender, and Ethnicity

The political discourse on social media is increasingly characterized by hate speech, which affects not only the reputation of individual politicians but also the functioning of society at large. In this work, we empirically analyze how the amount of hate speech in replies to posts from politicians on Twitter depends on personal characteristics, such as their party affiliation, gender, and ethnicity. We find that tweets are particularly likely to receive hate speech in replies if they are authored by (i) persons of color from the Democratic party, (ii) white Republicans, and (iii) women. Furthermore, our analysis reveals that more negative sentiment (in the source tweet) is associated with more hate speech (in replies). However, the association varies across parties: negative sentiment attracts more hate speech for Democrats (vs. Republicans). Altogether, our empirical findings imply significant differences in how politicians are treated on social media depending on their party affiliation, gender, and ethnicity.
Paper at WWW
Featured Research: Moralized language predicts hate speech on social media

This study provides large-scale observational evidence that moralized language fosters the proliferation of hate speech on social media. Specifically, we analyzed three datasets from Twitter covering three domains (politics, news media, and activism) and found that the presence of moralized language in source posts was a robust and meaningful predictor of hate speech in the corresponding replies. These findings offer new insights into the mechanisms underlying the proliferation of hate speech on social media and may help to inform educational applications, counterspeech strategies, and automated methods for hate speech detection.
- Research paper at PNAS Nexus (open access)
- Media coverage in Psychology Today
- Blog post on PsyPost
Featured Research: Negativity Drives Online News Consumption

Online media is important for society in informing and shaping opinions, hence raising the question of what drives online news consumption. Here, we analyze the causal effect of negative and emotional words on news consumption using a large online dataset of viral news stories. Specifically, we conducted our analyses using a series of randomized controlled trials (N = 22,743). Our dataset comprises ∼105,000 different variations of news stories from Upworthy.com that generated ∼5.7 million clicks across more than 370 million overall impressions. Although positive words were slightly more prevalent than negative words, we found that negative words in news headlines increased consumption rates (and positive words decreased consumption rates). For a headline of average length, each additional negative word increased the click-through rate by 2.3% Our results contribute to a better understanding of why users engage with online media.
- Research paper at Nature Human Behaviour
- Media coverage (selection): ARD, FAZ, Heise, Deutschlandfunk, ORF, Psychology Today, The Atlantic
Featured Research: Community-Based Fact-Checking Reduces the Spread of Misleading Posts on Social Media

Social media platforms increasingly rely on community-based fact-checking systems such as X’s Community Notes to combat misinformation at scale. In this study, we analyze more than 431 million reposts across 237,180 fact-checked cascades and provide large-scale causal evidence that community notes reduce the subsequent spread of misleading posts by 61.2% on average. We further find that community notes increase the likelihood that users delete misleading posts by 94.3%. However, notes often appear too late to prevent the early, most viral stage of diffusion, limiting their overall system-wide impact. Our findings highlight both the promise and current limitations of community-based fact-checking systems in reducing misinformation on social media.
- Research paper at Nature Communications (open access)
- Interview with FAZ
- Media Coverage (selection): The Washington Post, The Atlantic, TIME Magazine, ABC News, Poynter, BBC




