"AI is Only as Good as the Data" Interview with Gereon Rahnfeld
Gereon Rahnfeld is the team lead for Research Data at the German Informatics Society. In this role, he is responsible for NFDIxCS, the National Research Data Infrastructure for and with Computer Science. In this interview, he explains what NFDIxCS does, why research data management has become a decisive issue for AI and how current practices – from metadata work to peer review – are impacted.
Interview: Elena Müller, GI
NFDIxCS & The White Paper Behind This Conversation
Elena Müller: What is NFDIxCS and what does it aim to do?
Gereon Rahnfeld: NFDIxCS is the National Research Data Infrastructure for and with Computer Science. Our goal is to build an infrastructure that makes research data FAIR: findable, accessible, interoperable and reusable.
A recurring problem is that data may be stored somewhere, for example on GitHub, but cannot be re-executed. The software, the operating system or the precise execution environment are missing. Our answer is the Research Data Management Container, or RDMC. It stores a data point together with its software and execution environment so that experiments can be rerun and results reliably reused. That container is the core service NFDIxCS provides.
EM: How did the white paper come about, and what were you hoping to achieve with it?
GR: The white paper began as a Flex Fund subproject within NFDIxCS. Flex Funds allow NFDI consortia to support focused subprojects. Our team at the German Informatics Society had become increasingly interested in the intersectionof AI and research data management. We wanted to understand how AI is already being used in RDM, where the benefitsand risks lie, and how people in the NFDI context experience this in their everyday work.
The project was initiated by Ariana Dongus, a critical AI scholar. She designed the study, conducted six interviews with people in the NFDI environment and wrote the first draft. When she left the GI, I took over, revised and expanded the text and finalised it with our team. The paper is scheduled for formal publication in mid to lateAugust 2026, and we will present its findings at conferences such as KI 2026 and IJCAI-ECAI 2026.
Data as the Foundation of AI
EM: AI research is unusually dependent on data, yet reproducibility is a persistent problem. What keeps researchers from practising solid research data management?
GR: From interviews we conducted for a white paper on AI in research data management, three barriers came up repeatedly.
First, research data management costs time and effort. Researchers are already under pressure to publish, so RDM is often treated as an additional chore rather than as part of the core work.
Second, there is a lack of concrete know-how. Many people understand in principle that data should be reusable, but they are unsure what exactly to document, which standards to use or where to deposit data. RDM is still poorly anchoredin training and curricula.
Third, incentives are weak. Almost everyone agrees that research data management is “a good thing”, but it rarely plays a decisive role in project evaluation or career progression. If it does not matter for funding or jobs, it tends to slide down the priority list.
EM: Why is this particularly critical for the AI community?
GR: Data is foundational for AI: you need it to train, validate and benchmark models. The cliché that AI is only as goodas the data we feed it holds true. If the underlying data is poorly documented, inconsistently processed or biased, the systems built on top of it will reflect that.
In AI, responsible data management is not a luxury; it is a precondition for reproducible and robust systems. If you cannot tell which version of a dataset was used, which preprocessing steps were applied or whether test data leaked into the training set, then claimed advances in performance become hard to evaluate. That undermines not only individual projects, but the credibility of the field as a whole.
EM: How does NFDIxCS aim to help with these practical challenges?
GR: We work on two levels:
On the technical side, the RDMC makes it possible to keep data together with the exact software stack and execution environment needed to run it. In fields like AI, where reproducing experiments often depends on subtle environment details—library versions, GPU drivers, container configurations—this is crucial. It turns archived data into something that can be executed again, not just a snapshot that looks complete but cannot be usedin practice.
At the same time, NFDIxCS also sees itself as a place where communities meet. We bring together people from AI and from research data management—for example by taking part in conferences such as KI 2026 and HCAI and by involvingAI researchers in our activities. The idea is to align expectations: what AI researchers need in terms of data and reproducibility, and what RDM experts can realistically provide in terms of infrastructure, standards and support.
Metadata and the Validation Burden
EM: When you spoke to researchers for the white paper, what did you learn about how AI is already used in research data management itself?
GR: Two points were particularly striking.
The first is that metadata generation and enrichment is currently the central use case for AI in research data management. All interviewees mentioned it. If data is to be findable and reusable, you need good descriptions, keywordsand provenance information—what the dataset contains, how it was created, which versions exist, how it relates to other resources. Yet this work is often seen as something that happens, if at all, after the “real” research is finished.
AI is increasingly used to support or partially automate this step. For example, systems can suggest keywords, extract entities, generate draft descriptions or map free-text information to controlled vocabularies. Interviewees see clear value in this, because metadata work is both crucial and time-consuming. For AI research, this is especially important: rich and consistent metadata makes it easier to understand which datasets were used for training, how benchmark sets evolved and where possible sources of bias or contamination might lie.
The second point is what we call the validation burden. When AI generates metadata or other research-relevant output, you cannot simply accept it without checking. Someone has to verify that the result is correct and sensible, because these systems can hallucinate, misclassify or overlook important nuances. You introduce AI to save time, but then you still need human verification. That tension between the promise of efficiency and the need for oversight runs through many of the examples we heard.
For AI-focused projects, this tension is sharp. On the one hand, there is a real need to handle growing volumes of dataand documentation. On the other, if you outsource too much of the work to opaque tools and skip careful checking, you create new risks—incorrect metadata, misleading summaries, or undocumented transformations that make later interpretation and reuse difficult.
EM: Beyond metadata and validation, the white paper also looks at peer reviewing processes. What concerns you there?
GR: It is likely that AI will become more present in peer review, and that worries me. Peer review has long been understood as a human mechanism of quality control. If AI systems become heavily involved in both writing and reviewing manuscripts, the process risks losing the core of what it was supposed to be: an informedhuman judgement about the quality and originality of a piece of work.
At the same time, this situation forces us to reflect on our practices. We may have to ask whether the familiar model of submitting papers and reviewing them as we do now is still the best way to organise scientific communication and quality assurance. It could open debates about alternative models—more open forms of review, more emphasis on sharing data and code alongside articles—that are better suited to a research landscape in which such tools are ubiquitous.
Looking Ahead
EM: Your white paper concludes with several recommendations. If you had to highlight one for the AI community, which would it be?
GR: It would be the development of shared guidance for documenting the use of AI in data and software workflows. Using AI where it is obviously helpful—for example in metadata generation—will probably happen anyway. Agreeing on how we document its role is a genuine community task.
For AI research, this is crucial. If a model, a dataset or a benchmark has been shaped by AI-assisted tools at various stages, that should be transparent. Without shared standards for such documentation, it becomes difficult to understand, compare or reproduce AI-supported research. You end up with systems whose provenance is opaque, even ifthe code and data are nominally available.
EM: Finally, where do you want to take this work next?
GR: On the one hand, I want to examine more closely the specific tensions we identified, such as the validation burden, and see how researchers deal with them in practice. On the other, I want to broaden the range of voices. For this first white paper we mainly spoke to people within the NFDI context. In future, I want to include researchers and practitioners who work with AI and research data management outside that environment as well. For us, the white paper is the beginning of a longer conversation, not the final word.
EM: Thank you very much for the conversation.
GR: Thank you.
Transcribed and revised with Otter.ai
