About Experience Publications News Projects

Derck Prinzhorn

Derck Prinzhorn

AI Safety & Security

About

I'm a Member of Technical Staff at Exponential Security Labs, where I work on AI safety and security for agentic systems.

My research focuses on AI control, automated attacks, monitoring, and oversight of increasingly capable models. I recently completed my MSc in Artificial Intelligence at the University of Amsterdam, with thesis research at the Max Planck Institute for Intelligent Systems and ELLIS Institute Tübingen.

I also run Prinzhorn Solutions, helping organizations manage the risks of adopting AI systems. Previously, I worked as a Research Engineer at Aithos, as AI Architect at the Dutch Police, and co-founded Wisr. I graduated cum laude from my bachelor's, receiving the Amsterdam AI Thesis Award for my work on uncertainty quantification.

Experience

Exponential Security Labs

Aug 2026 – present
Member of Technical Staff

Building self-improving red-teaming and guardrail agents to secure agentic systems.

Max Planck Institute

Jan 2026 – Aug 2026
Research Intern

Research on AI control under supervision of Maksym Andriushchenko.

Prinzhorn Solutions

Apr 2025 – present
Founder

Helping companies understand and manage risks associated with adopting AI systems.

Aithos

Apr 2025 – Jan 2026
Research Engineer

Worked on evals for AI value systems and moral competence.

Wisr

Sep 2024 – Oct 2025
Co-Founder

Worked on a startup helping teachers save time with AI grading.

University of Amsterdam

Jan 2024 – Feb 2025
Research Intern

Conformal prediction for time series; 3D diffusion models for radiotherapy dose prediction; physics benchmarking in video generation models.

Dutch Police

Apr 2023 – Apr 2025
AI Architect

Defined reference architectures for AI, MLOps, and AI security.

Education

MSc Artificial Intelligence University of Amsterdam, 2023 – 2026
BSc Artificial Intelligence University of Amsterdam, 2020 – 2023

Publications

ResearchArena Figure 1: task setting, red-team agent, submission scoring, and blue-team monitoring

arXiv 2026

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

L. Libon, B. Rank, J. Yeon, D. Schmotz, J. Qin, D. Donnelly, D. Prinzhorn, M. Andriushchenko

Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems

ICML 2026

Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems

C. Zhang, D. Cherniavskii, [...], D. Prinzhorn, et al.

Conformal TSD

COPA 2024

Conformal Time Series Decomposition with Component-wise Exchangeability

D. Prinzhorn, T. Nijdam, P. van der Linden, A. Timans

FairAC

TMLR 2024

Reproducibility Study of FairAC

G. de Jong, M. Meijer, D. Prinzhorn, H. Ruiter

Highlights / News

2026
Aug

Joined Exponential Security Labs as Member of Technical Staff.

Jul

Released ResearchArena, a framework for evaluating sabotage and monitoring in automated AI R&D.

Mar

Participated in the Apart Research AI Control Hackathon, producing Controlling the Researcher.

Jan

Started thesis research on AI control at MPI-IS and ELLIS Institute Tübingen.

2025
Jul

HIVE paper accepted at the Beyond Euclidean Workshop, ICCV 2025.

Apr

Joined Aithos for pluralistic alignment research. Founded Prinzhorn Solutions.

2024
Dec

Presented FairAC reproduction as a poster at NeurIPS 2024.

Sep

Co-founded Wisr (EdTech). Presented conformal time series paper at COPA in Milan.

Jul

Joined SPAR, working on AI control with Aryan Bhatt from Redwood Research.

May

FairAC reproducibility paper accepted at TMLR. Conformal decomposition paper accepted at COPA.

2023
Nov

Received the AmsterdamAI thesis award for uncertainty quantification in time series.

Sep

Started MSc in Artificial Intelligence at the University of Amsterdam.

Projects

AI Control Evaluation Structure

Mar 2026

Controlling the Researcher: AI Control Evaluations for Automated AI R&D

Control evaluations for AI agents doing ML research. Subtle sabotage embedded in artifacts evades nearly all monitors, while obvious side tasks are reliably caught.

HIVE

Oct 2025

HIVE: Hyperbolic Visualization Explorer

Interactive dashboard for exploring hyperbolic embeddings with curvature-aware projections and multiple interaction modes. Published at ICCV 2025.

Injecting Image Guidance

Jul 2025

Injecting Image Guidance into Diffusion Models

Guide Stable Diffusion with both text and a reference image at inference time, without retraining. A lightweight aligner bridges the image-text embedding gap.

Radiotherapy Dose Prediction Architecture

Sep 2024

Transformer-Based Radiotherapy Dose Prediction

Deep learning for radiotherapy dose prediction in head and neck cancer. Extended UNETR with physics-informed losses and a sequential RNN decoder (~10% improvement).

View all projects