Showing posts with label Quality Assessment. Show all posts
Showing posts with label Quality Assessment. Show all posts

Tuesday, April 5, 2016

Garbage in, Garbage Out: Data Collection, Quality Assessment and Reporting Standards for Social Media Data Use in Health Research, Infodemiology and Digital Disease Detection

Background
Social media have transformed the communications landscape. People increasingly obtain news and health information online and via social media. Social media platforms also serve as novel sources of rich observational data for health research (including infodemiology, infoveillance, and digital disease detection detection). While the number of studies using social data is growing rapidly, very few of these studies transparently outline their methods for collecting, filtering, and reporting those data. Keywords and search filters applied to social data form the lens through which researchers may observe what and how people communicate about a given topic. Without a properly focused lens, research conclusions may be biased or misleading. Standards of reporting data sources and quality are needed so that data scientists and consumers of social media research can evaluate and compare methods and findings across studies.

Objective
We aimed to develop and apply a framework of social media data collection and quality assessment and to propose a reporting standard, which researchers and reviewers may use to evaluate and compare the quality of social data across studies.

Methods
We propose a conceptual framework consisting of three major steps in collecting social media data: develop, apply, and validate search filters. This framework is based on two criteria: retrieval precision (how much of retrieved data is relevant) and retrieval recall (how much of the relevant data is retrieved). We then discuss two conditions that estimation of retrieval precision and recall rely on—accurate human coding and full data collection—and how to calculate these statistics in cases that deviate from the two ideal conditions. We then apply the framework on a real-world example using approximately 4 million tobacco-related tweets collected from the Twitter firehose.

Results
We developed and applied a search filter to retrieve e-cigarette–related tweets from the archive based on three keyword categories: devices, brands, and behavior. The search filter retrieved 82,205 e-cigarette–related tweets from the archive and was validated. Retrieval precision was calculated above 95% in all cases. Retrieval recall was 86% assuming ideal conditions (no human coding errors and full data collection), 75% when unretrieved messages could not be archived, 86% assuming no false negative errors by coders, and 93% allowing both false negative and false positive errors by human coders.

Conclusions
This paper sets forth a conceptual framework for the filtering and quality evaluation of social data that addresses several common challenges and moves toward establishing a standard of reporting social data. Researchers should clearly delineate data sources, how data were accessed and collected, and the search filter building process and how retrieval precision and recall were calculated. The proposed framework can be adapted to other public social media platforms.

Below:  The archive (a+b+c+d), retrieved tweets (a+b), and relevant tweets (a+c+e) in Twitterverse



Below:  The average limits of 95% confidence intervals for recall (vertical axis) as the sample size of unretrieved messages increases (horizontal axis), fixing the sample size of retrieved data at 3000



Full article at:   http://goo.gl/WtdPlj

By:  1Health Media Collaboratory, Institute for Health Research and Policy, University of Illinois at Chicago, Chicago, IL, United States
Yoonsang Kim, Health Media Collaboratory, Institute for Health Research and Policy, University of Illinois at Chicago, Westside Research Office Building, M/C 275, 1747 W Roosevelt Rd, Chicago, IL, 60608, United States, Phone: 1 312 413 7596, Fax: 1 312 996 2703




Saturday, December 19, 2015

Quality of HIV Testing Data Before and After the Implementation of a National Data Quality Assessment and Feedback System

CONTEXT:
In 2010, the Centers for Disease Control and Prevention (CDC) implemented a national data quality assessment and feedback system for CDC-funded HIV testing program data.

OBJECTIVE:
Our objective was to analyze data quality before and after feedback.

DESIGN:
Coinciding with required quarterly data submissions to CDC, each health department received data quality feedback reports and a call with CDC to discuss the reports. Data from 2008 to 2011 were analyzed.

SETTING:
Fifty-nine state and local health departments that were funded for comprehensive HIV prevention services.

PARTICIPANTS:
Data collected by a service provider in conjunction with a client receiving HIV testing.

INTERVENTION:
National data quality assessment and feedback system.

MAIN OUTCOME MEASURES:
Before and after intervention implementation, quality was assessed through the number of new test records reported and the percentage of data values that were neither missing nor invalid. Generalized estimating equations were used to assess the effect of feedback in improving the completeness of variables.

RESULTS:
Data were included from 44 health departments. The average number of new records per submission period increased from 197 907 before feedback implementation to 497 753 afterward. Completeness was high before and after feedback for race/ethnicity (99.3% vs 99.3%), current test results (99.1% vs 99.7%), prior testing and results (97.4% vs 97.7%), and receipt of results (91.4% vs 91.2%). Completeness improved for HIV risk (83.6% vs 89.5%), linkage to HIV care (56.0% vs 64.0%), referral to HIV partner services (58.9% vs 62.8%), and referral to HIV prevention services (55.3% vs 63.9%). Calls as part of feedback were associated with improved completeness for HIV risk (adjusted odds ratio [AOR] = 2.28; 95% confidence interval [CI], 1.75-2.96), linkage to HIV care (AOR = 1.60; 95% CI, 1.31-1.96), referral to HIV partner services (AOR = 1.73; 95% CI, 1.43-2.09), and referral to HIV prevention services (AOR = 1.74; 95% CI, 1.43-2.10).

CONCLUSIONS:
Feedback contributed to increased data quality. CDC and health departments should continue monitoring the data and implement measures to improve variables of low completeness.

Purchase full article at:   http://goo.gl/GCCXdW

By:   Beltrami J1, Wang G, Usman HR, Lin L.
  • 1US Public Health Service and Division of HIV/AIDS Prevention, Centers for Disease Control and Prevention, Atlanta, Georgia (Dr Beltrami and Mr Wang); Alberta Health Services, Surveillance and Reporting, Edmonton, Alberta, Canada (Dr Usman); and Department of Mathematical Sciences, Montana State University, Bozeman (Dr Lin).