# Manjit Pokhrel > AI Safety Researcher researching multilingual LLM alignment failures and whether hardware/systems-level optimizations (quantization, compilation, power draw) degrade or leak that safety behavior. Based in Kathmandu, Nepal. - **Name**: Manjit Pokhrel - **Tagline**: AI Safety Researcher — Multilingual Alignment & Systems Security - **Website**: https://manjitpokhrel.com.np - **Location**: Kathmandu, Nepal - **Primary Affiliations**: - Independent AI Safety Researcher (2025 – Present) - Research Intern — LLM Development at ILPRL (Information and Language Processing Research Lab), Kathmandu University (2026 – Present) - Contributor at MLCommons (DMLR & AI Risk and Reliability Working Groups, 2026 – Present) - Undergraduate, BSc Computer Science at Kathmandu University (2025 – 2029) - **Contact**: manjitpokhrel42@gmail.com - **Google Scholar**: https://scholar.google.com/citations?user=NaxYesYAAAAJ&hl=en - **GitHub**: https://github.com/manjitpokhrel - **LinkedIn**: https://linkedin.com/in/manjitpokhrel - **X (Twitter)**: https://x.com/manjitpokhrel_ ## Core Research Areas 1. **Multilingual LLM Alignment & Jailbreak Asymmetry**: - Examining safety divergence between high-resource languages (English) and low-resource/regional languages (Nepali, Hindi, code-switched Devanagari/Latin). - Attack vectors exploiting sub-tokenization splits, morpho-syntactic boundaries, and cultural framing. 2. **Mechanistic Interpretability & Model Internals**: - Residual stream activation extraction, directional ablation, and safety preservation under model transformations. 3. **Systems & Hardware Security (Compilers & Quantization)**: - Investigating whether post-training optimizations (torch.compile, Triton kernels, TensorRT, 4-bit NF4/INT4 quantization) erode model refusal directions or leak confidential safety guardrails. 4. **Inference Efficiency & Hardware-Aware AI**: - Activation sparsity (GhostWeight CUDA kernels), GPU kernel optimization, and peak inference energy dynamics on seasonal renewable electrical grids. ## Key Publications & Formal Acknowledgments - **Lost in Translation: Safety Alignment Failures in Nepali and Code-Switched Variants of Instruction-Tuned Large Language Models** - Authors: Manjit Pokhrel - Year: 2026 - Publisher: Zenodo - DOI: [10.5281/zenodo.19764520](https://doi.org/10.5281/zenodo.19764520) - Code: [github.com/manjitpokhrel/NASB-Nepali-Safety](https://github.com/manjitpokhrel/NASB-Nepali-Safety) - Abstract: Establishes a 73.7% safety-bypass rate in Nepali vs. 0.0% in English across frontier instruction-tuned LLMs (Gemini, Llama-3, Qwen, Gemma). Introduces Vajra Morphing, an intra-sentential multi-script attack exploiting sub-tokenization gaps. - **Quantifying AI Inference as an Emerging Load on Nepal's Seasonal Hydropower Grid (NASB-E)** - Authors: Manjit Pokhrel - Conference: ReSET 2026 (Kathmandu University) - Focus: Peak vs. average energy consumption analysis for LLM inference refusal vs. bypass on developing country renewable grids. - **Google AI Vulnerability Reward Program (VRP)**: - Research: Multilingual AI safety-alignment bypass research. Status: Triaged (2026). - **Meta Whitehat**: - Research: Multilingual safety asymmetry across model tiers. Status: Acknowledged (2026). - **GhostWeight: Training-Free Activation Sparsity for LLM Inference on Consumer Hardware**: - Package: PyPI ([pypi.org/project/ghostweight/](https://pypi.org/project/ghostweight/)) - Repository: [github.com/manjitpokhrel/GhostWeight](https://github.com/manjitpokhrel/GhostWeight) ## Selected Projects - **NASB (Nepali Adversarial Safety Benchmark)**: - URL: https://manjitpokhrel.com.np/projects/nasb - Description: Benchmark of 385 structured adversarial queries across harm taxonomies with 1,200+ probes across frontier models. - **The Eleventh Optimization**: - URL: https://manjitpokhrel.com.np/projects/eleventh-optimization - Description: Mechanistic interpretability study evaluating refusal direction ablation through torch.compile, Triton, and quantization layers. - **GhostWeight**: - URL: https://manjitpokhrel.com.np/projects/ghostweight - Description: Custom CUDA kernels skipping inactive MLP neurons (~27.3% sparsity in Qwen2.5-7B) yielding 110.5% inference speedup with zero measured perplexity loss. - **Quantifying AI Inference on Nepal's Seasonal Hydropower Grid**: - URL: https://manjitpokhrel.com.np/projects/nasb-energy - **Nepali LLM Fine-Tuning Pipeline**: - URL: https://manjitpokhrel.com.np/projects/nepali-llm-finetuning - Description: Low-resource parameter-efficient fine-tuning (PEFT/LoRA) pipeline for Devanagari script NLP. - **MiniGPT**: - URL: https://manjitpokhrel.com.np/projects/minigpt - Description: 211K parameter transformer written in pure NumPy with manual backpropagation (no autograd). ## Full Details - Complete research digest for LLM grounding: https://manjitpokhrel.com.np/llms-full.txt - Comprehensive Curriculum Vitae: https://manjitpokhrel.com.np/cv - Articles & Essays: https://manjitpokhrel.com.np/blogs