About the project
Position in the Digital Futures research matrix
This project sits within the Trust research theme, which addresses the foundations of safe, reliable, and accountable digital systems, and connects to the Smart Society societal context, which looks at how digital technologies shape public life and institutions. By working at this intersection, the project bridges technical AI safety research with the societal and human dimensions of AI misuse.
Objective
The project sets out to understand how generative AI systems can be exploited for socially harmful purposes, through jailbreaks, adversarial prompting, and other unsafe forms of human-AI interaction. Rather than treating this as a narrow cybersecurity issue, the project looks at broader harms such as manipulation, coercive persuasion, misinformation, and intimate surveillance. Its goals are to map emerging misuse scenarios, identify research gaps across AI safety, HCI, and sociotechnical security, build interdisciplinary collaborations, explore ethical approaches to red-teaming, and produce an initial roadmap for future research in this space.
Background
The project responds to a growing body of evidence that safety-aligned AI systems can still be manipulated into producing harmful or restricted outputs, despite existing safeguards. Techniques for bypassing these safeguards continue to evolve and circulate within online communities, often outpacing efforts to detect and prevent them. While most existing AI safety research concentrates on technical robustness and content filtering, far less attention has gone to socially situated harms like coercive control, deceptive influence, and interpersonal abuse enabled by AI. This project aims to help close that gap between technical jailbreak research and human-centred research on digital harm.
Cross-disciplinary collaboration
The project brings together expertise from human-centred cybersecurity, HCI, control systems security, and software supply chain security, combined through a shared program of exploratory research and four interdisciplinary workshops covering AI-enabled interpersonal harm, cybersecurity misuse, misinformation and manipulative communication, and ethical red-teaming methodologies.
PIs
- Asreen Rostami (PI) – Senior Researcher, RISE Research Institutes of Sweden
- Nicolas Harrand (Co-PI) – Senior Lecturer, Stockholm University
- Henrik Sandberg (Co-PI) – Professor, KTH Royal Institute of Technology

