Integrating Answer Set Programming and Large Language Models for Reliable Structured Knowledge Extraction
Main Article Content
Abstract
Answer Set Programming (ASP) and Large Language Models (LLMs) offer complementary strengths in symbolic reasoning and natural language understanding. In this work, we present an integrated framework for reliable structured knowledge extraction from natural language that combines the syntactic capabilities of LLMs with the semantic constraints provided by ASP. Structured representations in the form of ASP facts are obtained by guiding LLM generation through domain specific prompt templates and background knowledge specified in a declarative YAML configuration, and by validating and completing the extracted information using an ASP knowledge base.
Building on this framework, we introduce grammar constrained decoding as a principled mechanism to enforce structural and semantic alignment between LLM outputs and target symbolic representations. By leveraging formal grammars, we gain fine grained control over the generation process, substantially reducing hallucinations and limiting unnecessary verbosity, while preserving extraction quality. We study three grammar formats commonly used to represent structured knowledge, namely CSV, Datalog, and JSON, and analyze their impact on correctness, efficiency, and computational cost.
We evaluate the proposed approach on benchmarks derived from ASP Competitions. The results show that the integration of ASP consistently improves extraction accuracy over vanilla LLMs, particularly for smaller models. Moreover, grammar constrained decoding maintains these gains while significantly improving generation efficiency. Among the considered formats, CSV provides the best trade-off between accuracy and cost, JSON achieves the highest extraction quality at the expense of increased verbosity, and Datalog exhibits lower robustness in the extraction process. Overall, the results demonstrate that combining ASP with grammar constrained LLM output yields a reliable, scalable, and cost-effective approach to structured knowledge extraction.