Skip to main navigation Skip to search Skip to main content

SpannerLib: Embedding Declarative Information Extraction in an ImperativeWorkflow

  • Dean Light
  • , Ahmad Aiashy
  • , Mahmoud Diab
  • , Daniel Nachmias
  • , Stijn Vansummeren
  • , Benny Kimelfeld

Research output: Contribution to journalConference articlepeer-review

Abstract

Document spanners have been proposed as a formal framework for declarative Information Extraction (IE) from text, following IE products from the industry and academia. Over the past decade, the framework has been studied thoroughly in terms of expressive power, complexity, and the ability to naturally combine text analysis with relational querying. This demonstration presents Spanner- Lib—a library for embedding document spanners in Python code. SpannerLib facilitates the development of IE programs by providing an implementation of Spannerlog (Datalog-based document spanners) that interacts with the Python code in two directions: rules can be embedded inside Python, and they can invoke custom Python code (e.g., calls to ML-based NLP models) via user-defined functions. The demonstration scenarios showcase IE programs, with increasing levels of complexity, within Jupyter Notebook.

Original languageEnglish GB
Pages (from-to)4281-4284
Number of pages4
JournalProceedings of the VLDB Endowment
Volume17
Issue number12
DOIs
StatePublished - 2024
Event50th International Conference on Very Large Data Bases, VLDB 2024 - Guangzhou, China
Duration: 24 Aug 202429 Aug 2024

ASJC Scopus subject areas

  • Computer Science (miscellaneous)
  • General Computer Science

Fingerprint

Dive into the research topics of 'SpannerLib: Embedding Declarative Information Extraction in an ImperativeWorkflow'. Together they form a unique fingerprint.

Cite this