DatasetRadar — Dataset Intelligence Platform
End-to-end system monitoring news & journal sites to automatically detect cited public datasets via self-hosted LLMs, vector deduplication (pgvector), and a confidence-gated human review queue.
Building production-grade data pipelines, full-stack data products, and stateful agentic AI systems.
Experienced in taking initiatives from technical architecture through to deployed, containerized solutions — with hands-on expertise across pgvector semantic search, LangGraph stateful multi-actor agent workflows, and autonomous pipeline operations.
End-to-end system monitoring news & journal sites to automatically detect cited public datasets via self-hosted LLMs, vector deduplication (pgvector), and a confidence-gated human review queue.
Modular Python monorepo for engineering and deploying stateful, multi-actor agent workflows: dynamic tool-calling, conversational session checkpointers (MemorySaver), and self-healing SQL guardrails.
Architecting full-stack data products, applied AI/LLM engineering (LangChain, LangGraph, RAG, pgvector), automated ETL pipelines, and internal validation suites powering the Dataful public datasets platform.
I take a project from a written spec and architecture doc all the way through to a fully working system with a real API, database, and frontend — not just prototypes. DatasetRadar is my proof: I built a complete, running data-intelligence platform with pgvector-backed semantic search, a self-hosted LLM detection pipeline, a confidence-gated human review UI, and both dev and production Docker deployments — not a partial build, a shipped product.
I build semantic search, retrieval-augmented Q&A, embedding-based deduplication, and stateful multi-actor agent workflows using LangChain and LangGraph — I layer AI onto real systems and structured monorepos rather than bolt it on as a standalone demo.
I design reusable, composable AI skills and stateful graph architectures (such as my LangGraph Playground Monorepo) that encode an entire team’s conventions — directory structure, coding standards, validation rules, cyclic tool-calling loops, and checkpoint-backed persistent memory — into self-correcting workflows, complete with safety scaffolding like branch checks, query guardrails, and step-by-step rollback. This is a distinct capability from writing pipelines by hand: I build the tooling that lets a team, or an agent, reliably reproduce my standards without me in the room.
I build systems that run unattended on a schedule, touch real infrastructure — chat, git, version control — and get hardened over multiple iterations based on real failures. This goes a step beyond scripting: I design safeguards, like branch checks and reply-detection fallbacks, so an autonomous process doesn’t do the wrong thing, and knows when to stop and ask for help.
That’s my core data-engineer persona: scraping, data quality profiling, catching subtle pipeline bugs, and packaging tools so others can run them without me in the room.