PGxMine: Text mining for curation of PharmGKB

Pac Symp Biocomput. 2020:25:611-622.


Precision medicine tailors treatment to individuals personal data including differences in their genome. The Pharmacogenomics Knowledgebase (PharmGKB) provides highly curated information on the effect of genetic variation on drug response and side effects for a wide range of drugs. PharmGKB's scientific curators triage, review and annotate a large number of papers each year but the task is challenging. We present the PGxMine resource, a text-mined resource of pharmacogenomic associations from all accessible published literature to assist in the curation of PharmGKB. We developed a supervised machine learning pipeline to extract associations between a variant (DNA and protein changes, star alleles and dbSNP identifiers) and a chemical. PGxMine covers 452 chemicals and 2,426 variants and contains 19,930 mentions of pharmacogenomic associations across 7,170 papers. An evaluation by PharmGKB curators found that 57 of the top 100 associations not found in PharmGKB led to 83 curatable papers and a further 24 associations would likely lead to curatable papers through citations. The results can be viewed at and code can be downloaded at

Publication types

  • Research Support, N.I.H., Extramural

MeSH terms

  • Computational Biology
  • Data Mining / methods
  • Databases, Genetic
  • Humans
  • Knowledge Bases
  • Pharmacogenetics*
  • Precision Medicine* / methods