Improving the Utility of Tobacco-Related Problem List Entries Using Natural Language Processing

Daniel R Harris; Darren W Henderson; Alexandria Corbeau

Improving the Utility of Tobacco-Related Problem List Entries Using Natural Language Processing

AMIA Annu Symp Proc. 2021 Jan 25:2020:534-543. eCollection 2020.

Authors

Daniel R Harris^{1

2}, Darren W Henderson², Alexandria Corbeau²

Affiliations

¹ Institute for Pharmaceutical Outcomes and Policy, College of Pharmacy, University of Kentucky, Lexington, Kentucky 40506.
² Center for Clinical and Translational Sciences, University of Kentucky, Lexington, KY 40506.

PMID: 33936427
PMCID: PMC8075422

Abstract

We present findings on using natural language processing to classify tobacco-related entries from problem lists found within patient's electronic health records. Problem lists describe health-related issues recorded during a patient's medical visit; these problems are typically followed up upon during subsequent visits and are updated for relevance or accuracy. The mechanics of problem lists vary across different electronic health record systems. In general, they either manifest as pre-generated generic problems that may be selected from a master list or as text boxes where a healthcare professional may enter free text describing the problem. Using commonly-available natural language processing tools, we classified tobacco-related problems into three classes: active-user, former-user, and non-user; we further demonstrate that rule-based post-processing may significantly increase precision in identifying these classes (+32%, +22%, +35% respectively). We used these classes to generate tobacco time-spans that reconstruct a patient's tobacco-use history and better support secondary data analysis. We bundle this as an open-source toolkit with flow visualizations indicating how patient tobacco-related behavior changes longitudinally, which can also capture and visualize contradicting information such as smokers being flagged as having never smoked.

Publication types

Research Support, N.I.H., Extramural

MeSH terms

Electronic Health Records*
Humans
Medical Records, Problem-Oriented / standards*
Natural Language Processing*
Nicotiana
Tobacco Use / adverse effects*

Grants and funding

UL1 TR001998/TR/NCATS NIH HHS/United States