Oudoum Ali Houmed
AI Safety & Security
Currently: Research Engineering Intern at Kappa Santé — Generative Modeling & AI Robustness · Healthcare
About
I'm an M.Sc. student in Data, Knowledge and Hybrid Artificial Intelligence (DKAI) at Paris-Saclay University and an AI Safety Research Fellow at Black in AI Safety & Ethics under Krystal Jackson (UC Berkeley, Center for Long-Term Cybersecurity). My fellowship project is “Can We Trust Deception Monitors for AI Agents? A Robustness-Gap Protocol and the Limits of Adaptive Evasion.”
My fellowship work asks how much deception detection survives as adversary budget rises: a robustness-gap protocol for monitors on AI agents (surface classifiers, residual-stream probes, and CoT controls), with a retention gate so a broken agent is not scored as evasion. Broader interests include adversarial robustness, red teaming, mechanistic interpretability, and control mechanisms for frontier models.
Alongside the fellowship, I work as a Research Engineering Intern at Kappa Santé on constraint-guided virtual patient generation for clinical datasets. Previously, I was a Research Assistant at ANÖROM, where I studied the adversarial robustness of medical diagnostic models. I have completed the ARENA curriculum, OxML 2025, and BlueDot Impact's Technical AI Safety course.
Experience
Kappa Santé
May 2026 – Sep 2026Applied research on constraint-guided virtual patient generation for longitudinal clinical datasets: conditional generation pipelines with explicit filtering for statistical fidelity and temporal coherence, and evaluation of downstream model safety with a focus on calibration and generalization under distribution shift.
Black in AI Safety & Ethics
May 2026 – Jul 2026Fellowship project under Krystal Jackson (UC Berkeley CLTC): Can We Trust Deception Monitors for AI Agents? A Robustness-Gap Protocol and the Limits of Adaptive Evasion — on Llama-3.3-70B with the published Apollo probe, an adversary-budget ladder (b0–b4), and a retention gate before reporting detection drop. Findings feed technical standards, safety controls, and policy recommendations. Completed the intensive 5-week ARENA curriculum.
ANÖROM — Big Data & AI Dept
Oct 2024 – Jun 2025Investigated the vulnerability of medical diagnostic architectures to adversarial perturbations using gradient-based optimization. Implemented white-box attacks (FGSM, PGD) to quantify the robustness of tumor detection models; contributed empirical findings toward a forthcoming manuscript (Adversarial Threats to Safety-Critical Medical AI).
ADEO Cybersecurity
Feb 2024 – Jun 2024Achieved an internship grade of 95%. Detected and analyzed RDP brute-force attacks using Wazuh SIEM; performed deep-packet inspection with Wireshark and web-app pen-testing with Burp Suite; conducted malware scanning, static/dynamic reverse engineering, and forensic analysis with MDR tools.
AI Security & Defence Lab (AISEC LAB)
Jul 2023 – Sep 2023Co-developed Cyber Inspector, a WAF-integrated ML system for malicious query detection. Managed the full project lifecycle: data collection, preprocessing, model training (SVM, RF), and deployment (Detection of Malicious Web Queries with Machine Learning).
Research Stack
Libraries & Frameworks
Core Concepts
General
Publications
Manuscript · ANÖROM Big Data & AI · 2025
Adversarial Threats to Safety-Critical Medical AI: A Security Assessment of 20 Deep Learning Tumour Detectors in Brain MRI and Kidney CT
Certificates
Highlights / News
Selected for the AI Safety Research Fellowship at Black in AI Safety & Ethics (Alignment & Security track), supervised by Krystal Jackson (UC Berkeley CLTC). Started Can We Trust Deception Monitors for AI Agents?, measuring how monitors hold up under adaptive adversaries.
Joined Kappa Santé in Paris as a Research Engineering Intern, working on constraint-guided virtual patient generation.
Formalized the Narrative Susceptibility Score (NSS) in the LLM Biases for Political Conspiracies project, with a 600+ prompt evaluation protocol.
Started the M.Sc. in Data, Knowledge and Hybrid Artificial Intelligence (DKAI) at Paris-Saclay University.
First paper published in Artificial Intelligence Studies: “A Systematic Review on Evolution of Endpoint Security” (with O. Ceran).
Attended OxML 2025 (Oxford Machine Learning Summer School) and completed BlueDot Impact's Technical AI Safety course.
Joined ANÖROM (Big Data & AI Dept) in Ankara as a Research Assistant on AI robustness.
Cybersecurity Specialist Intern at ADEO Cybersecurity (Istanbul): Wazuh SIEM, Wireshark, Burp Suite, and MDR forensics (internship grade 95%).
Graduated in the top 10% of the B.Sc. Computer Engineering cohort at Necmettin Erbakan University, on a full scholarship.
Summer Research Intern at AISEC LAB (Konya): co-developed Cyber Inspector, a WAF-integrated ML system for malicious query detection.
Completed the UNDP-IICPSD Machine Learning Bootcamp.
Projects
2026 · Alignment & Security
Can We Trust Deception Monitors for AI Agents? A Robustness-Gap Protocol and the Limits of Adaptive Evasion
Fellowship project (Black in AI Safety & Ethics): a robustness-gap protocol for deception monitors on AI agents — Apollo residual-stream probe, surface and CoT controls, adversary-budget ladder b0–b4, and a retention gate before reporting Δdet.
2026 · NLP & AI Safety
LLM Biases for Political Conspiracies — Constrained Prompt Evaluation
Controlled Entity × Narrative × Template evaluation of political
conspiracy-framed prompts across 20 leaders, 10 topics, and 3 styles
(Direct / Creative / Socratic) — 600 prompts scored into a
Narrative Susceptibility Score (NSS). Compares Mistral, GPT-4o Mini, DeepSeek, and
Llama 3.1 70B on safety, hallucination type, and entity-level fairness
(run 20260208_200540_385119c8).