Display Settings:

Format

Send to:

Choose Destination

    AMIA Annu Symp Proc. 2005:156-60.

    Empirical data on corpus design and usage in biomedical natural language processing.

    Cohen KB, Fox L, Ogren PV, Hunter L.

    Center for Computational Pharmacology, U. of Colorado School of Medicine, USA. kevin.cohen@gmail.com

    This paper describes the design of six publicly available biomedical corpora. We then present usage data for the six corpora. We show that corpora that are carefully annotated with respect to structural and linguistic characteristics and that are distributed in standard formats are more widely used than corpora that are not. These findings have implications for the design of the next generation of biomedical corpora.

    PMID: 16779021 [PubMed - indexed for MEDLINE]

    PMCID: 1560643

    Supplemental Content

    Click here to read Click here to read