AI Safety & Alignment · Governance

Subramanyam Sahoo

Email

Portrait of Subramanyam Sahoo

Alignment fails in practice before it fails in theory. I work on catching it.

My research spans reinforcement-learning post-training, mechanistic interpretability, the science of evaluations, and governance frameworks that hold when institutions come under strain, built over 2.5+ years across academic and independent settings.

Post-training RL Mechanistic Interpretability Science of Evaluations AI Agents AI Governance

Research across Erdős AI Lab · Cambridge AI Safety Hub · EleutherAI · UC Berkeley (BASIS) · Apart Research · Open Democracy Institute · Martian · InsightX Lab (Northeastern University)

Research

I study the failure modes that surface only when systems meet the real world: reward hacking under distribution shift, sycophancy and specification gaming, silent failures in latent reasoning, and the institutional mechanisms that catch what technical safeguards miss.

It has produced first-author workshop papers at ICLR, NeurIPS, AAAI, ICML and NAACL, including an oral presentation at the ICLR 2026 AI for Peace Workshop and joint work with the InsightX Lab at TrustNLP, and contributions to two EvalEval Coalition studies published at the ICML 2026 main conference. The complete, current record lives in one place: Google Scholar. Longer-form essays and research notes are on Substack.

Positions & Fellowships

2026 – present
Research Collaborator
InsightX Lab, Khoury College of Computer Sciences, Northeastern University

NLP interpretability with Dr. Divya Chaudhary's group; joint paper at TrustNLP, NAACL 2026.

Jun – Aug 2026
Summer Research Resident

Mechanistic interpretability as a lens for understanding model deviation from sanctioned behaviour.

Apr – Jun 2026
CORDA Democracy Fellow
Open Democracy Institute

Integrity disclosures for generative AI in democratic information environments.

Dec 2025 – Feb 2026
MARS 4.0 — Mentorship for Alignment Research Students
Cambridge AI Safety Hub

Alignment research on RL agents under the mentorship of Prof. Fernando Rosas (University of Sussex), working towards ICLR 2027.

Aug – Dec 2025
Researcher, Summer of Open AI Research (E-SOAR)
EleutherAI

Led a project on prompt optimisation for verifiable hallucination reduction.

Apr 2025 – present
Independent Contractor
Outlier AI

Designing synthetic datasets for RL-style post-training and evaluation under controlled task distributions.

Feb – Oct 2025
AI Policy Fellow
Berkeley AI Safety Initiative (BASIS), UC Berkeley — remote

Governance research for the Berkeley AI Safety Initiative.

Aug 2022 – Jul 2024
Teaching Assistant & Lab Supervisor
NIT Hamirpur

Taught and supervised AI, machine learning, deep learning and systems courses and labs for B.Tech and dual-degree students; developed curriculum and assessment materials; organised the AMRIT-2023 and MINDS-2023 conferences.

2018 – 2019
Earlier

Summer internships: sentiment-analysis pipeline at LIT, Bhubaneswar (2019); database indexing systems at Naresh i Technologies, Hyderabad (2018).

Education

2022 – 2024
M.Tech, Computer Science & Engineering (Artificial Intelligence)
National Institute of Technology, Hamirpur

CGPA 9.38/10 · Gold Medalist, Summa Cum Laude · Batch topper.

2016 – 2020
B.Tech with Honours, Computer Science & Engineering
Parala Maharaja Engineering College

CGPA 8.68/10 · Ranked in the top 5 of the college.

Hackathons & Competitions

Aug 2026
Jul 2026
Jun 2026
Global South AI Safety Hackathon · Apart Research
Apr 2026
Mar 2026
Feb 2026
The Technical AI Governance Challenge · Apart Research
Sep 2025
CBRN AI Risks Research Sprint · Apart Research
Apr 2025
Berkeley AI Policy Hackathon
UC Berkeley

Project later accepted as a NeurIPS 2025 workshop paper.

Research Grants

Nov 2025 – Feb 2026
Martian — Research Grant · USD 6,000
Principal Investigator

Mechanistic interpretability: experimental design, model analysis and dissemination.

Oct 2025
Apart Research — Research Grant · USD 500

Pilot experiments and preliminary analyses for an independent research contribution.

Recognition

Apr 2026
Y Combinator Startup School
Bangalore, India
Summer 2025
Harvard Technical AI Safety Fellowship & Harvard AI Policy Fellowship
Dual fellowships awarded
2025
Apart Lab Studio Internship

Accepted on the strength of a mechanistic-interpretability hackathon project. Continuing collaboration with Amirali Abdullah (Martian).

2025
Carnegie Mellon University — Generative AI & LLM certificate programme

Selected with a partial scholarship; withdrew due to funding constraints.

2024
Full fee waiver — “Harms and Risks of AI in the Military” workshop
Mila — Quebec AI Institute, Montreal
2024
Climate Change AI Summer School
Mila — Quebec AI Institute
2024
ACM India Summer School — Responsible and Safe AI
IIT Madras
2024
ACM India Summer School — Generative AI for Text
IIT Gandhinagar

Media

Oct 2024
Invited speaker on AI safety, Odisha AI Conference 2024
Virtual event hosted in the USA

Service & Volunteering

2026 – present
Reviewer — AI safety workshops at A* conferences

Workshop reviewing at ICLR, ICML and ACL, and conference reviewing for COLM.

Jan – Apr 2026
AI Safety Camp — 11th Edition
Jul – Oct 2025
Mentor, Paragon AI Policy Fellowship
May – Jun 2025
Mentor, LatinX AI Club — 2025 Edition
Nov 2024 – May 2025
AI Safety Camp — 10th Edition

Selected Programmes

Aug – Sep 2026
CLR SPI Fundamentals Program
Center on Long-Term Risk
Mar – May 2026
Cooperative AI Fundamentals
Cooperative AI Foundation
May – Sep 2025
AI Agents and Law
Vista Institute for AI Policy

Beyond Research

Languages: English (full professional), Odia (native), Hindi (limited working), Sanskrit (limited working), Spanish (elementary).

Off hours: critically acclaimed podcasts.

Work With Me

I'm open to research contractor roles, PhD opportunities, and collaborations in alignment, evaluations and AI governance. If your work touches any of the problems above, the fastest route is email.

Email me