BORDEA, Georgeta; FARALLI, S; MOUGIN, Fleur; BUITELAAR, P; DIALLO, Abdourahmane Gayo

Metadatos

Mostrar el registro completo del ítem

Licencia de uso del documento

BORDEA, Georgeta

Bordeaux population health [BPH]

FARALLI, S

MOUGIN, Fleur

Bordeaux population health [BPH]

Idioma

Communication dans un congrès avec actes

Este ítem está publicado en

Proceedings of The 12th Language Resources and Evaluation Conference, Proceedings of The 12th Language Resources and Evaluation Conference, 2020-05, Marseille. 2020p. 2341–2347

European Language Resources Association

Resumen en inglés

In this work, we address the task of extracting application-specific taxonomies from the category hierarchy of Wikipedia. Previous work on pruning the Wikipedia knowledge graph relied on silver standard taxonomies which can only be automatically extracted for a small subset of domains rooted in relatively focused nodes, placed at an intermediate level in the knowledge graphs. In this work, we propose an iterative methodology to extract an application-specific gold standard dataset from a knowledge graph and an evaluation framework to comparatively assess the quality of noisy automatically extracted taxonomies. We employ an existing state of the art algorithm in an iterative manner and we propose several sampling strategies to reduce the amount of manual work needed for evaluation. A first gold standard dataset is released to the research community for this task along with a companion evaluation framework. This dataset addresses a real-world application from the medical domain, namely the extraction of food-drug and herb-drug interactions.< Leer menos

Palabras clave

ERIAS

URI

https://oskar-bordeaux.fr/handle/20.500.12278/25809

Centros de investigación

Bordeaux Population Health Research Center (BPH) - UMR 1219