EvalFlow
Backend-first eval platform for LLM apps with prompt versioning, benchmark runs, LLM-as-judge scoring, and prompt promotion.
Architecting intelligent systems from chaos. Specialized in Machine Learning and deep-diving into MLOps.
Backend-first eval platform for LLM apps with prompt versioning, benchmark runs, LLM-as-judge scoring, and prompt promotion.
Autonomous AI code review agent for pull requests, focused on bug detection, regression checks, and security-oriented reasoning.
Agentic video generation system for prompt-driven scene planning, retrieval, orchestration, and automated production pipelines.
End-to-end NLP and MLOps pipeline with experiment tracking, monitoring, artifact versioning, and production telemetry for model visibility.
Rebuilt the core transformer architecture to study attention mechanics, encoder-decoder flow, and training behavior from scratch.
Implemented the original convolutional pipeline to understand early vision architectures, parameter sharing, and classifier construction.