Bhargava Sai Deepak Mathukumalli — Full-stack Applied-AI & Core Systems Engineer · Bengaluru, India

I ship full-stack AI products in production — and understand every layer beneath them.

Software engineer (SDE-2) working end to end: LLM agents and their evals, Next.js dashboards, Go/gRPC services — and, on my own time, the layer underneath it all, rebuilt from first principles: inference servers, trained models, CUDA kernels.

SDE-2 · ~2.5 yrs full-time production full-stack AI · Next.js · Go · FastAPI LLM agents · evals · MCP

In production — professional work

Shipped, full-stack, and used by customers.

In production Amnic · full-stack

Prism by Amnic

Built end to end — frontend, backend, services & infra

A pre-deploy cost-intelligence SaaS: it estimates the cloud-cost impact of infrastructure-as-code changes before they ship, and exposes it to AI clients through a Model Context Protocol server so Claude or Cursor can query scans in natural language.

Next.js · TypeScript FastAPI · GraphQL Go · gRPC MCP server Celery · Postgres Fly.io
In production Amnic · applied AI + evals

Amnic AI — agent & eval harness

LLM agent plus the framework that keeps it honest

A production LLM agent that answers cloud-cost questions over customer data via tool calls — and the LLM-as-judge evaluation framework I built to guard it: acceptance-criteria and golden-reference scoring, plus regression and investigation modes that replay historical tool-call traces.

LLM agent · tool-use LLM-as-judge evals regression + golden refs Anthropic · OpenAI · Bedrock

Company work — described at a high level, no code shared.

Built from scratch — the depth

The same stack, rebuilt from first principles.

L2 · The Modelabstraction

scribe

A ~30M-parameter GPT trained from scratch on TinyStories — kept GPT-2-compatible on purpose, so my own inference server can load it and stream stories.

val loss 1.74
~3.4h on MPS
TinyStories · GPT-2 compat
L1 · The Serversystems

ember

A mini-vLLM inference server: paged KV cache, continuous batching and an OpenAI-compatible streaming API — with GPT-2 itself rebuilt from scratch inside it.

paged KV cache
continuous batching
13 paging tests
L0 · The Metalhardware

cuda-kernels

Hand-written CUDA, escalating in difficulty: vector-add → tiled GEMM → fused softmax → FlashAttention with online softmax — benchmarked against cuBLAS and PyTorch.

FlashAttention
online softmax
vs cuBLAS / torch

Selected work — breadth

Beyond the ML stack.

quorum
Distributed Systems · Go

A distributed key-value store with a hand-written Raft consensus layer — leader election, log replication and snapshots — over an LSM engine, verified under Go's race detector.

athena-cost-guard
Developer Tooling · Python

A pre-flight cost estimator for AWS Athena with a @cost_guard budget decorator (sqlglot + Glue + S3) — published to PyPI.

text-diffusion
Generative ML · PyTorch

A tiny text-conditioned diffusion model — a mini Stable Diffusion — built from scratch and trained on Apple-silicon MPS.

kalman-timeseries
Numerical Methods · NumPy

An adaptive, robust Kalman filter in pure NumPy, written up as an academic-style LaTeX paper; the self-tuning variant cut error ~51% over the baseline.

clickstream-sessionizer
Data Engineering · PySpark

A medallion lakehouse (bronze → silver → gold) with gap-based sessionization, skew-salting, broadcast joins and Structured Streaming — validated on ~400k events.

All repositories on GitHub ↗

About

Shipping in production. Rebuilding the hard parts by hand.

Bhargava Sai Deepak

I'm a Software Engineer (SDE-2) with roughly 2.5 years of full-time experience, based in Bengaluru and currently at Amnic, where I work across the whole stack — a Next.js dashboard, Python and Go services, an LLM agent and the eval framework that keeps it correct.

The through-line: I don't like treating layers as magic. So on my own time I rebuild the ones I depend on — inference servers, consensus protocols, GPU kernels, trained models — from first principles, because reimplementing something is the only way I fully trust that I understand it. Those projects are open source: real code, real training runs, real benchmarks.

ECE dual degree (B.Tech + M.Tech), IIT Kharagpur — Vision & Intelligent Systems specialization, CS minor.

Put together: I can ship a full-stack AI product and also reason about what happens inside it, down to the kernel. If you're hiring for applied AI, ML, or backend that values both — let's talk.

Résumé (PDF)