AI/ML & Software Engineer specializing in LLMs, RAG, Agentic AI, Edge AI, and MLOps. Currently building production AI systems from inference to deployment at MulticoreWare.
I build AI systems that ship - from low-level model optimization on edge hardware to full-stack AI platforms serving real customers in production. My engineering focus spans building end-to-end MLOps pipelines, designing robust RAG architectures (embedding, re-ranking, retrieval), executing fine-tuning orchestration, and deploying multi-tenant adapter inference engines. I care just as much about the infrastructure layer as the models themselves. Using tools like AWS, Kubernetes, and Docker, I architect scalable backends featuring Stripe-based billing, LLM-token-based autoscaling, secure RBAC authentication, and comprehensive Prometheus observability to ensure production-grade reliability.
On the systems programming side, I write high-performance C++ for custom inference servers, leveraging asynchronous multi-threading and optimized KV Caching (achieving up to 33.71 tok/s). My expertise in Edge AI allows me to push hardware boundaries using state-of-the-art quantization and calibration techniques like LiteRT and AIMET. I recently applied these principles to architect an end-to-end AI-enabled IDE for embedded systems - featuring comprehensive model-optimization pipelines that take neural networks from raw import to hardware-ready deployment, alongside custom tooling like model visualizers, latency estimators, and debug dashboards.
Outside of my core engineering roles, I actively research and implement Agentic AI workflows, focusing on frameworks like LangChain, ReAct Agents, and the Model Context Protocol (MCP). My recent personal projects include a dual-pipeline medical diagnostic assistant combining fixed RAG with dynamic reasoning agents, and an emotion-aware Spotify recommendation engine driven by custom neural networks and LSTM-based generation. I graduated as the Founder Chancellor's Gold Medalist for my B.Tech in AI & ML, and I am always excited to tackle complex engineering challenges or talk shop on modern AI architecture.
Additionally, I am a 3x published book chapter author, with peer-reviewed contributions in the fields of Artificial Intelligence and Machine Learning.
Dual-pipeline diagnostic assistant for medical-field Q/A, utilizing LangChain ReAct orchestration with Llama 3.3 70B and FAISS-based RAG with a fine-tuned FLAN-T5 model.
Emotion-based and custom ML algorithm-based song recommendation web application integrated with Spotify API. Built custom CNN model from scratch using TensorFlow and Keras.
Conversational AI for code assistance integrating a custom Transformer model in PyTorch. Generates executable Python code from plain English prompts via a Discord chatbot and React/Flask web app, with live compilation via Judge0.
Flask application utilizing TensorFlow Object Detection and EasyOCR to extract vehicular number plates. Stores live data in Google Sheets, visualized via a PowerBI dashboard, and features an OpenAI-powered chatbot.
Machine Learning web app utilizing the Twitter API to fetch and analyze real-time highway traffic data. Features predictive modeling using Random Forest, Logistic Regression, and Naive Bayes Classifiers.
MulticoreWare · 2026
Recognized for exceptional
ownership, technical proficiency, and a proactive problem-solving approach
leading to smooth, timely feature deliveries.
SRIHER, Chennai · 2025
Awarded for outstanding
academic performance and achieving the highest CGPA in the graduating
cohort.
SRIHER, Chennai · 2024
Won top honors for
presenting an emotion-based music recommendation engine integrated with the
Spotify API.
I'm always open to discussing modern AI systems, model optimization, or production AI infrastructure. Whether you have a question or just want to say hi, my inbox is always open!