Derck Prinzhorn
AI Safety & Security
About
I'm a Member of Technical Staff at Exponential Security Labs, where I work on AI safety and security for agentic systems.
My research focuses on AI control, automated attacks, monitoring, and oversight of increasingly capable models. I recently completed my MSc in Artificial Intelligence at the University of Amsterdam, with thesis research at the Max Planck Institute for Intelligent Systems and ELLIS Institute Tübingen.
I also run Prinzhorn Solutions, helping organizations manage the risks of adopting AI systems. Previously, I worked as a Research Engineer at Aithos, as AI Architect at the Dutch Police, and co-founded Wisr. I graduated cum laude from my bachelor's, receiving the Amsterdam AI Thesis Award for my work on uncertainty quantification.
Experience
Exponential Security Labs
Aug 2026 – presentBuilding self-improving red-teaming and guardrail agents to secure agentic systems.
Max Planck Institute
Jan 2026 – Aug 2026Research on AI control under supervision of Maksym Andriushchenko.
Prinzhorn Solutions
Apr 2025 – presentHelping companies understand and manage risks associated with adopting AI systems.
Aithos
Apr 2025 – Jan 2026Worked on evals for AI value systems and moral competence.
Wisr
Sep 2024 – Oct 2025Worked on a startup helping teachers save time with AI grading.
University of Amsterdam
Jan 2024 – Feb 2025Conformal prediction for time series; 3D diffusion models for radiotherapy dose prediction; physics benchmarking in video generation models.
Dutch Police
Apr 2023 – Apr 2025Defined reference architectures for AI, MLOps, and AI security.
Publications
Highlights / News
Joined Exponential Security Labs as Member of Technical Staff.
Released ResearchArena, a framework for evaluating sabotage and monitoring in automated AI R&D.
Participated in the Apart Research AI Control Hackathon, producing Controlling the Researcher.
Started thesis research on AI control at MPI-IS and ELLIS Institute Tübingen.
HIVE paper accepted at the Beyond Euclidean Workshop, ICCV 2025.
Joined Aithos for pluralistic alignment research. Founded Prinzhorn Solutions.
Presented FairAC reproduction as a poster at NeurIPS 2024.
Co-founded Wisr (EdTech). Presented conformal time series paper at COPA in Milan.
Joined SPAR, working on AI control with Aryan Bhatt from Redwood Research.
FairAC reproducibility paper accepted at TMLR. Conformal decomposition paper accepted at COPA.
Received the AmsterdamAI thesis award for uncertainty quantification in time series.
Started MSc in Artificial Intelligence at the University of Amsterdam.
Projects
Mar 2026
Controlling the Researcher: AI Control Evaluations for Automated AI R&D
Control evaluations for AI agents doing ML research. Subtle sabotage embedded in artifacts evades nearly all monitors, while obvious side tasks are reliably caught.
Oct 2025
HIVE: Hyperbolic Visualization Explorer
Interactive dashboard for exploring hyperbolic embeddings with curvature-aware projections and multiple interaction modes. Published at ICCV 2025.
Jul 2025
Injecting Image Guidance into Diffusion Models
Guide Stable Diffusion with both text and a reference image at inference time, without retraining. A lightweight aligner bridges the image-text embedding gap.
Sep 2024
Transformer-Based Radiotherapy Dose Prediction
Deep learning for radiotherapy dose prediction in head and neck cancer. Extended UNETR with physics-informed losses and a sequential RNN decoder (~10% improvement).