About
I'm a machine learning researcher from Chennai, currently an ML intern at Glance and starting an MSc in Machine Learning at UCL.
My research is on AI safety and mechanistic interpretability. I want to understand what's actually happening inside language models, and use that understanding to build systems that stay safe by default.
I'm Tamil, from Chennai, and I follow Liverpool FC obsessively.
Education
Publications
Conversational Hallucination Drift: An Episodic Retrieval Framework and Error Taxonomy for Long-Term Memory Evaluation
Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures
Probing Reasoning Flaws and Safety Hierarchies with Chain-of-Thought Difference Amplification
Indian Grammatical Tradition-Inspired Universal Semantic Representation Bank (USR Bank 1.0)
Writing
Exploring Generalization in NLA's
Reproduction and follow-on experiments on Anthropic's neural network activation work, focused on training architecture and generalization as they relate to steering.
LessWrong · Jun 2026
Calibrating Activation Vectors using Norm
Activation steering usually hand-tunes injection magnitude. This post looks at calibrating vector norms directly instead of applying a uniform scale across layers.
LessWrong · Jun 2026
Experience
Designed an LLM-as-judge evaluation pipeline for catalog recommendations, testing prompting strategies (chain-of-thought, structured rubrics, multi-step verification) to improve judgment reliability and reduce false positives.
Developed a method to distill human annotator and LLM-judge signals into a cross-encoder, enabling recommendation scoring for zero-usage / cold-start users with no behavioral history.
Built and optimized catalog retrieval pipelines, improving content discovery across millions of items; achieved a 2× improvement in retrieval metrics by reimplementing SOTA embedding techniques from recent papers.
Awarded monetary prizes for top-tier performance in high-stakes predictive modeling and algorithmic problem-solving challenges.
Developed and optimized ML models under strict competitive constraints in a fast-paced environment.
Working with Ihor Kendiukhov to evaluate scalability and security guarantees of novel control protocol classes for advanced AI systems.
Extending Greenblatt et al.'s framework by designing hierarchical and parallel control topologies to optimize safety-usefulness trade-offs.
Conducting scaling experiments to verify generalization of control guarantees across diverse model capabilities.
Conducted independent research on mechanistic interpretability, with direct technical guidance from a Research Scientist at the UK AI Safety Institute.
Developed a framework to permanently embed steering vectors into model weights, enabling persistent safety behaviors without inference-time compute overhead.
Researched Universal Semantic Representations and developed techniques for generating coherent natural language sentences from abstract syntactic-semantic structures.
Worked on Controlled Image-to-Text Generation systems for scientific images to ensure accurate, context-aware, domain-specific textual descriptions.
Enhanced query performance by 25% by optimizing document embedding strategies for semantic similarity retrieval.
Automated mapping of legal clause dependencies, improving analysis efficiency by 30%.
Research Projects
ManifoldSteer: Geometry-Preserving Activation Steering in LLM Residual Streams
Geometry-preserving steering operators for LLM residual streams, built on an empirical layer-selection criterion derived from mapping Qwen2.5-3B's semantic submanifolds.
Mapped the Qwen2.5-3B residual stream to identify ~5–6D semantic submanifolds, establishing an empirical layer-selection criterion for safe steering.
Black-Scholes Modelling with Kolmogorov Arnold Networks
Applying KANs to financial derivatives pricing, exploring whether learnable activation functions improve on classical Black-Scholes approximations.
Working to improve the 72% accuracy achieved on the Synthetic European dataset by refining KAN architecture.
Grants
Lambda Labs Research Grant ($1,000 (scalable to $5,000))
Awarded for research on interpretability and routing dynamics of Mixture-of-Experts (MoE) models during inference. Work involves mathematical analysis of expert load balancing, probing memorization vs. generalization through expert usage patterns, and inference-only diagnostic tools.
Talks & Seminars
Academic Service
Reviewer · ICML 2026, AI-MATH Workshop
Reviewer · ICML 2026, Mechanistic Interpretability Workshop
Reviewer · ICML 2026, Failure Modes in Agentic AI Workshop
Reviewer · ICML 2026, SCALE Workshop
Reviewer · NeurIPS 2025, AI-MATH Workshop
Reviewer · AAAI 2026, Workshop on Shaping Responsible Synthetic Data in the Era of Foundation Models