Source
NLDB
DATE OF PUBLICATION
07/04/2025
Authors
Mikhail Salnikov Andrey Sakhovskiy Irina Nikishina Aida Usmanova Angelie Kraft Cedric Möller Debayan Banerjee Junbo Huang Longquan Jiang Rana Abdullah Xi Yan Elena Tutubalina Ricardo Usbeck Alexander Panchenko
Share

ShortPathQA: A Dataset for Controllable Fusion of Large Language Models with Knowledge Graphs

Abstract

In this work, we release the Shortest Path subgraph Question Answering(ShortPathQA) dataset, the first dataset that provides textual questions withpre-computed relevant subgraphs retrieved from the Wikidata knowledge graph,standardizing the evaluation framework for Knowledge Graph Question Answering(KGQA). For this purpose, we utilize the Mintaka dataset for both trainingand testing and additionally create a manual question-answering subset for testing.Our baseline experiments with both supervised approaches and unsupervisedLarge Language Model (LLM) inference indicate that even a simplified KGQAformulation with given Knowledge Graph (KG) subgraphs and candidate answersremains challenging. Our analysis has shown that LLMs are unable to correctlyprocess and utilize graph data structures without detailed prompt engineering ormodel tuning. This limitation highlights the need for the creation of this dataset asa training ground for the development of methods that enable LLMs to work moreeffectively with graph data.


Join AIRI