Senior Data & AI Systems Engineer

Venu Madhav Sambarapu

Building production-grade data pipelines, full-stack data products, and stateful agentic AI systems.

Experienced in taking initiatives from technical architecture through to deployed, containerized solutions — with hands-on expertise across pgvector semantic search, LangGraph stateful multi-actor agent workflows, and autonomous pipeline operations.

4+ Yrs IT(+11 Acad.)
Senior Data & AI Engineer
Full-Stack
FastAPI · pgvector · React
LangGraph
Stateful Multi-Actor Agents
Autonomous
Supervised ETL Triage

Featured Flagship Projects

All Projects →
Full-Stack Data Product

DatasetRadar — Dataset Intelligence Platform

End-to-end system monitoring news & journal sites to automatically detect cited public datasets via self-hosted LLMs, vector deduplication (pgvector), and a confidence-gated human review queue.

FastAPIpgvectorPostgreSQLReact 18CeleryOllama
Stateful AI Agents

LangGraph Playground Monorepo

Modular Python monorepo for engineering and deploying stateful, multi-actor agent workflows: dynamic tool-calling, conversational session checkpointers (MemorySaver), and self-healing SQL guardrails.

LangGraphLangChainSQLiteMemorySaverPydantic

Have questions about my technical background?

My interactive portfolio assistant knows my project specs, ETL architectures, and agent implementations.

Chat with Portfolio Assistant ✦

Experience & Education

Senior Data Engineer
4+ Years IT Experience
Factly Media and Research · Dec '21 – Present · Hyderabad, IN

Architecting full-stack data products, applied AI/LLM engineering (LangChain, LangGraph, RAG, pgvector), automated ETL pipelines, and internal validation suites powering the Dataful public datasets platform.

Prior Academic ExperienceAssistant Professor · Various Engineering Institutions · Aug '10 – Nov '21 (11+ Years) — Undergraduate engineering pedagogy in digital electronics, communication systems, and data analytics; NAAC accreditation analytics and Python workshops.

Education & Certifications

Post Graduate Certification in Data Science
IIIT Bangalore & upGrad · Aug '20 – Mar '21 · Secured 90%
M.Tech, Digital Electronics & Communication Systems
Mahaveer Institute of Science and Technology · Jan '12 – Jan '15

End-to-End Data Product Delivery

I take a project from a written spec and architecture doc all the way through to a fully working system with a real API, database, and frontend — not just prototypes. DatasetRadar is my proof: I built a complete, running data-intelligence platform with pgvector-backed semantic search, a self-hosted LLM detection pipeline, a confidence-gated human review UI, and both dev and production Docker deployments — not a partial build, a shipped product.

Applied AI/LLM Engineering

I build semantic search, retrieval-augmented Q&A, embedding-based deduplication, and stateful multi-actor agent workflows using LangChain and LangGraph — I layer AI onto real systems and structured monorepos rather than bolt it on as a standalone demo.

Agent/Skill Engineering

I design reusable, composable AI skills and stateful graph architectures (such as my LangGraph Playground Monorepo) that encode an entire team’s conventions — directory structure, coding standards, validation rules, cyclic tool-calling loops, and checkpoint-backed persistent memory — into self-correcting workflows, complete with safety scaffolding like branch checks, query guardrails, and step-by-step rollback. This is a distinct capability from writing pipelines by hand: I build the tooling that lets a team, or an agent, reliably reproduce my standards without me in the room.

Production Automation & Agentic Ops

I build systems that run unattended on a schedule, touch real infrastructure — chat, git, version control — and get hardened over multiple iterations based on real failures. This goes a step beyond scripting: I design safeguards, like branch checks and reply-detection fallbacks, so an autonomous process doesn’t do the wrong thing, and knows when to stop and ask for help.

Core Data Engineering Fundamentals

That’s my core data-engineer persona: scraping, data quality profiling, catching subtle pipeline bugs, and packaging tools so others can run them without me in the room.