Mogankumar Narsozhan
AI/ML Engineer

Building and deploying state-of-the-art ML systems at scale.
Focused on LLMs, efficient neural architectures, and production MLOps.

Advancing the frontier of applied ML

I'm an ML Engineer / Backend with 2 years shipping production systems solo, from distributed ingestion pipelines to LLM-backed inference infrastructure. My work spans high-throughput async pipelines, semantic search, vector databases, and RAG systems, all deployed at scale on AWS.

Previously a Founding ML Engineer at Kubby, where I owned the full backend lifecycle end-to-end — architecture, deployment, and optimization. Before that, at Hewlett Packard Enterprise working on enterprise NFV infrastructure.

I hold an M.S. in Computer Science (AI Track) from SUNY Buffalo and a B.Tech in Electrical and Electronics Engineering from NIT Puducherry. I specialize in turning complex data into scalable, impactful AI solutions for real-world challenges.

Machine Learning

Training, fine-tuning, and deployment of predictive ML systems and neural networks.

Generative AI & RAG

Engineering RAG-based assistants, vector DB integration, and semantic search optimization.

MLOps & Systems

End-to-end pipelines, Docker, Openstack, and reducing end-to-end latency in production.

Software Engineering

Scalable system design, Python development, and translating unstructured data into insights.

Specialization & Skills

AI solutions for real-world problems. Also well-versed in Software Development and Linux.

Education

SUNY Buffalo

State University of New York at Buffalo, NY, USA

Masters in Computer Science (AI Track) | Aug, 2024 - Dec, 2025

GPA: 3.6/4

Courses: Introduction to Machine Learning, Data Intensive Computing, Algorithms Analysis and Design, Intro to Pattern Recognition, Deep learning, Operating systems, Data Modelling Query Language, Computer Vision

NIT Puducherry

National Institute of Technology, Puducherry, India

Bachelor of Technology, Electrical and Electronics Engineering | 2019 - 2023

GPA: 8.63/10

GitHub Projects

Reinforcement Learning

Deep Learning

Reinforcement Learning

Implemented reinforcement learning techniques such as SARSA and Q-Learning to build a firefighter simulation environment for dynamic problem-solving.

CODE

Virtual Mouse

Object Detection

Virtual Mouse

A Python-based Virtual Mouse that uses hand gestures for cursor control, clicking, scrolling, and taking screenshots. Powered by OpenCV, PyAutoGUI, and MediaPipe for a touch-free experience.

CODE

Bird Flock Simulation

Big Data Analytics

Bird Flock Simulation

Designed and implemented a bird flock simulation using PySpark to model complex behaviors in distributed systems, showcasing the scalability of big data processing frameworks.

CODE

Face Recognition System

Computer Vision

Face Recognition

Developed a facial recognition system using SVM and OpenCV for accurately identifying individuals from photographs. Included techniques for image preprocessing and feature extraction.

CODE

Neural Networks & CNN

Deep Learning

Neural Networks & CNN

This project focuses on building fully connected neural networks (NN) and convolutional neural networks (CNN).

CODE

Self-Healing Technical Documentation

Python

Self-Healing Documentation

A documentation drift detection and repair system that lives inside the CI/CD pipeline, solving the problem of temporal decay as code evolves continuously.

CODE

Semantic Caching Layer for LLM APIs

Python

Semantic Caching Layer

A middleware proxy service between applications and LLM providers that checks for semantically similar past queries before sending expensive API requests.

CODE

Autoencoders for Anomaly Detection

Jupyter Notebook

Anomaly Detection

Built an LSTM-based autoencoder for unsupervised anomaly detection on AWS EC2 CPU metrics, flagging anomalies via reconstruction error spikes.

CODE

Healthcare Simple RAG

Jupyter Notebook

Healthcare Simple RAG

Implemented a simple Retrieval-Augmented Generation (RAG) system tailored for healthcare applications to provide context-aware responses.

CODE

AI Sketch to UI Converter

Deep Learning

Sketch to UI

Converts hand-drawn UI sketches into functional HTML code. Integrates a custom YOLOv8 detector with a CNN + Attention + Transformer + GRU pipeline.

CODE

UrbanTwinAI

Streamlit

UrbanTwinAI

An interactive digital twin platform for generative city simulation. Visualizes how urban design changes impact traffic, heat distribution, and air quality using AI-driven simulations.

CODE

COVID-19 Classification

Deep Learning

COVID-19 Classification

Classifies chest X-ray images into Normal, COVID-19, and Viral Pneumonia using a two-layer CNN architecture, demonstrating high accuracy.

CODE

Work Experience

Kubby

Founding ML Engineer | San Francisco, CA (remote) | Feb 2026 – Present

  • Owned full end-to-end delivery of transformer-based predictive ML systems: scoped requirements, designed evaluation frameworks measuring accuracy, latency, and business performance criteria, and drove deployment to production, achieving sub-200ms inference latency across 50K+ product records.
  • Led structured model performance reviews with customers, directly influencing retention and expansion decisions.
  • Owned the full delivery lifecycle: problem framing, evaluation design, production deployment, and ongoing monitoring. Aligning customer engineering teams around reliable, observable outcomes.

Kubby

AI/ML Engineer | San Francisco, CA (remote) | May 2025 – Dec 2025

  • Engineered a production RAG-based AI shopping assistant (12+ live sources) delivering context-aware recommendations and semantic Q&A with sub-second response times.
  • Diagnosed a critical semantic search latency bottleneck through systematic performance analysis; engineered a Redis precomputation solution with vector database integration that reduced end-to-end pipeline latency by 75% (1.2s → 300ms), directly improving customer-facing product experience.
  • Built a transaction analytics engine translating unstructured behavioral signals into structured business intelligence dashboards, supporting executive and customer stakeholder decision-making.
  • Reduced data ingestion latency 75% via SQS + ECS re-architecture; documented decisions to align customer engineering teams around reliable, observable delivery.

Hewlett Packard Enterprise

SVC Info Developer | Full-time | August 2023 - May 2024

  • Assisted in creating RedHat Package Managers using Network Functions Virtualization Director (NFVD), HP Service Activator, and Workflow Designer and installed various versions, patch bundles, and hotfixes.
  • Executed the migration of Samsung EnterpriseDB to PostgresDB.
  • Experience: Linux Commands, Openstack, Docker, Networking (TCP/IP)

Hewlett Packard Enterprise

Cybersecurity Engineer | Internship | January - July 2023

  • Performed Breach and Attack Simulation (BAS): a proactive security approach that tests the organization's security posture and fixes vulnerabilities by simulating real-world attacks.
  • Monitored all the services, hosts, databases, and data centers, using NAGIOS XI, which are associated with a particular project, and resolved any issues that occurred
  • Experience: Hypervisors (Virtual Machines, VMWare), Docker, and Ubuntu.

Zebo.ai

ML Intern | Internship | October - November 2022

  • An intermediate level of Python was studied for the training, testing, and cross-validation of data followed by the features and labels pickling, scaling, techniques, and Error Metrics.
  • Implemented many Machine Learning algorithms.
  • Experience: Python, Machine Learning Algorithms, Data Structures.

Contact Me

If you'd like to collaborate or learn more about my work, feel free to reach out!