"We’re Deploying Systems Faster Than We Can Build Safeguards" Interview with Daniel Sonntag
Daniel Sonntag is scientific director at the German Research Center for Artificial Intelligence (DFKI), where he works on intelligent user interfaces, human-centered AI and machine learning, with a focus on real-world applications in domains such as medicine. Elena Müller spoke with him at this year’s IJCAI conference in Bremen. In this interview, he looks back on his long history with IJCAI, explains why symbolic methods will not disappear in the age of large language models, defends a positive notion of “weak AI”, and warns that highly capable systems are being deployed faster than society can establish safeguards and meaningful oversight.
Interview: Elena Müller, GI
Daniel Sonntag and IJCAI
Elena Müller: To start, could you introduce yourself and tell me why you’re here at IJCAI?
Daniel Sonntag: I’m Daniel Sonntag. I’m scientific director at DFKI, working on intelligent user interfaces, human-centered AI and machine learning.
EM: Have you been to IJCAI before, and what are you most looking forward to this year?
DS: I remember my first IJCAI well: 2007 in Hyderabad, where we took a rickshaw from the hotel to the conference resort. Ten years later I was in Melbourne, another very good event. Since then, we’ve presented several demo systems at IJCAI, and this year I’d like to see more human-centered AI represented in the demos and in the industrial track. It’s not only foundational work that deserves scientific recognition. Practical applications should be acknowledged at the major AI conferences too.
Hybrid AI: Symbolic Meets Data-Driven
EM: Rule-based AI was once considered a dead end, but hybrid systems that combine data-driven learning with older, rule-based approaches are back in fashion. Is AI coming full circle, or is this something new?
DS: Not full circle. We’re not returning to the old expert-systems paradigm, because data-driven models can do far more than they could ten years ago. But learning from data alone is often insufficient, and we need more transparency and controllability, especially of inference. So the question isn’t which paradigm wins, it’s how we combine their strengths in systems people can actually use and oversee.
In our own work on interactive machine learning, experts shape and correct model behavior in real time. Expert knowledge is itself a symbolic signal, which is why I don’t think we’ll ever get rid of symbolic systems entirely.
EM: How far can today’s data-driven models go on their own, without any rule-based help?
DS: Large language models will keep getting better at reasoning, but tasks that need explicit constraints and guarantees will still require symbolic methods. I’d argue for a division of labor between learned models and explicit expert knowledge and judgment that the system can draw on in real time when needed. That means the system is never as autonomous as it might appear.
The Case for Weak AI
EM: Should AI’s understanding of the world be learned purely from data, or does it need built-in rules and real-world constraints?
DS: I’m not convinced a single, better world model is the missing ingredient for useful AI. In medicine, an important application domain for us, we don’t need systems with a complete human-like understanding of the world. We need systems that are reliable within carefully designed tasks, like recognizing malignant tissue in an image, with their results checked and corrected by experts when necessary.
That’s why I remain an advocate of “weak AI” in a positive sense: systems with bounded responsibilities, embedded in human workflows such as clinical workflows. Those limits make them more transparent, since there’s an implicit regulation in not claiming autonomy beyond what an expert can validate. What matters is transparency, trustworthiness, interaction design, and, especially in medicine, human oversight.
EM: Which is the bigger unsolved problem: strict logical reasoning, or spotting patterns in messy, real-world information?
DS: The reasoning problems themselves haven’t changed much in the years I’ve worked in application domains like medicine. What I find most useful is common-sense and practical reasoning, and there we’ve had the same open problems for twenty years: uncertain, vague, incomplete information, and limited resources to compute or retrain.
It’s the same challenge the semantic web community has long faced: producing answers that stay cautious where evidence is thin, explain their assumptions, and revise conclusions when needed. I don’t see a way to solve common-sense reasoning under uncertainty without involving experts and end users directly, drawing on their domain knowledge to fill the gaps.
From Benchmarks to Socio-Technical Systems
EM: AI research tends to focus on individual systems and benchmark scores, but future systems will operate as participants in complex social settings alongside humans and other AI. Should the field shift toward evaluating AI as part of socio-technical systems?
DS: Yes, we’re often asking the wrong questions, focusing too much on benchmark scores and on intelligence in the abstract. From an application-oriented perspective, daily socio-technical practice should be central in fields such as medicine, sustainability, industry and home care: minimizing unacceptable risk and understanding societal impact. Conferences like IJCAI are still far from that practical ground.
A very accurate system can still fail if it doesn’t fit the workflow or if users can’t tell when to trust it. Workflow integration, accountability, safety and trust matter at least as much as raw performance.
Cyber Risks and the Case for Regulation
EM: Large language models can now generate code, plan complex tasks and adapt strategies from feedback. How much will frontier AI change the planning and execution of cyber attacks?
DS: I’m not a cybersecurity expert, but agentic AI, combining planning, automation and code generation, clearly lowers the threshold for harmful activity. That makes strong evaluation scenarios essential, and the AI community needs more time to build safeguards before highly autonomous systems are widely deployed.
It echoes early semantic-web-services ideas about automatic discovery, composition and execution of services: an old idea in a modern, much more complex form. Cyber attacks are a clear example of the kind of harmful execution such safeguards need to catch.
EM: Is the EU prepared for increasingly capable AI that can be used for cyber attacks and other harms, and is there promising regulation in Germany?
DS: It isn’t really a German problem, it’s geopolitical and societal. AI is no longer just IT. It’s becoming a geopolitical, socio-technical infrastructure that touches work, healthcare, security, public institutions and the distribution of power. My concern is that we’re deploying increasingly capable systems faster than we can build safeguards.
Regulation shouldn’t be seen only as a brake on innovation. Done well, it buys time to develop security measures, evaluation methods and the institutional capacity to test them. That’s close to Stuart Russell’s argument that we need more regulation and a real culture of safety before AI systems can be shown to be safe and beneficial.
Whether in clinical decision support or cybersecurity, we should be cautious of systems that make decisions without meaningful human control. Society doesn’t want systems to act completely autonomously without human oversight. That’s one way to define human-centered AI, even if the concept doesn’t translate directly to the context of cyber attacks.
EM: Thank you very much for the conversation.
DS: Thank you. I’m looking forward to all the sessions, demonstrations and industry pitches at this year’s IJCAI here in Bremen.
Transcribed and revised with Otter.ai
