
3
1
India
5+ years of experience
I build intelligent systems at the intersection of AI, infrastructure, and software engineering, with a focus on designing AI systems from the ground up from local AI runtimes and agent architectures to retrieval systems, orchestration layers, and AI-native developer tooling. My work spans embedded databases, developer infrastructure, full-stack engineering, and scalable system design, driven by a deep interest in low-level architectures, memory systems, and reliable intelligent software. Alongside engineering, I bring strong business thinking, product intuition, and a research-driven mindset, combining technical depth with the ability to identify opportunities, shape product direction, and turn ambitious ideas into practical systems. I enjoy rapidly experimenting with prototypes, solving complex problems creatively, and leading projects from concept to execution. Over the years, Iβve built AI tools and full-stack platforms, led technical initiatives and research efforts, collaborated across diverse domains, and won multiple hackathons and technical competitions. My long-term focus is on AI infrastructure, local-first AI, intelligent operating systems, multimodal systems, retrieval architectures, and developer tooling that shapes the next generation of AI products.

The Hybrid Token-Efficient Routing Agent (Rauto) solves the "Dynamic Model Problem" by intelligently orchestrating user queries across a configurable pool of local and cloud-based Large Language Models. Instead of defaulting to expensive frontier models for every task, Rauto employs a fast, local helper model to instantly classify query domains and dynamically assess difficulty. It powers its Cost-Aware Decision Engine using a mathematical Bayesian Belief Fusion framework. This engine merges static empirical capability priors (derived from industry benchmarks like MMLU and HumanEval) with dynamic evidence retrieved from a local vector-based Experience Memory database. If the system encounters a highly novel query, it intelligently falls back to a 3-Layer Graph Memory structure to map abstract reasoning requirements to the best-suited model. Furthermore, Rauto ensures maximum accuracy via a rigorous verification layer. Utilizing an AST Code Sandbox and a local LLM Judge with majority-vote heuristics, it autonomously self-corrects failures before returning an answer. The result is a highly scalable, Docker-containerized orchestration pipeline that slashes API costs without sacrificing output quality.
13 Jul 2026