Loading...
Loading...
Real-world impact: From 10K RPS architectures to products serving 51M+ users at Cafe Bazaar
The production AI agent ran on five separate contexts. That meant no persistent memory across sessions, no way to answer from documents a user uploads, and no visibility into what the agent was doing while it worked. Consolidating this into one unified system was a full architecture decision: migration strategy, scalability, and integration.
Led migration of the agent from five separate contexts into one unified architecture. As part of that migration, shipped a persistent per-user memory system, so context carries across sessions instead of resetting each time. Built a RAG pipeline that answers from documents a user uploads directly, with real citations. Added security guardrails that keep the user-facing agent out of its own prompts, source code, and infrastructure. Extended the agent with multimodal generation (image and video from a prompt) and live execution transparency, streaming plan and step status to the user as it works. Also owns prompt and context engineering and the eval suite (LLM-as-judge plus regression cases) that gates every agent change.
Growing an AI team from 2 to 11 engineers in a year, through a summer of war and power outages, needs engineers who can really build and reason about systems, not just talk about them well in an interview. That hiring standard has to stay consistent for every candidate instead of depending on personal impressions.
Designed a two-stage engineering interview process (a background and design-depth conversation, followed by a live coding and audit session) with an explicit, evidence-based scoring rubric to keep evaluations consistent and defensible. Personally screens CVs, runs interviews, and mentors new engineers through onboarding.
Design and implement a high-performance architecture capable of handling 10,000 requests per second while serving 51+ million users globally. The existing system was experiencing performance bottlenecks and needed a complete architectural overhaul.
Led the design and implementation of a distributed microservices architecture with strategic optimizations: implemented efficient caching layers with Redis, designed horizontal scaling strategies with Kubernetes, optimized database queries and indexing, and used gRPC for inter-service communication to reduce latency.
Critical services written in Python were facing scalability issues and performance bottlenecks. The system needed to handle increasing load while maintaining service quality for millions of users.
Led a strategic migration of key services from Python to Go, focusing on high-traffic services first. Built out test coverage for each migrated service, rolled out gradually behind feature flags, and profiled each service with gprof before and after cutover. Established best practices for Go development across the organization.
The system had evolved into an over-fragmented microservices architecture with inefficiently separated services. This created unnecessary network overhead, complex deployments, and difficult-to-maintain codebases.
Led the initiative to analyze service boundaries and consolidate inefficiently separated microservices. Identified services with tight coupling, merged them strategically, and redesigned APIs for better service boundaries. Implemented domain-driven design principles for clearer service definitions.
With rapid team growth and increasing system complexity, maintaining up-to-date documentation became a major challenge. Manual documentation was often outdated and developers spent significant time searching for information.
Developed an automated documentation generation system that extracted API documentation from code comments, generated interactive API references, and created system architecture diagrams. Integrated with CI/CD pipeline to ensure documentation was always current.
Build a culture where engineers take complete ownership of their services from development to production, including deployment, monitoring, and data-driven decision making.
Implemented "zero to 100" ownership model where engineers owned their entire stack: wrote services, managed CI/CD pipelines, self-managed Kubernetes deployments, dockerized applications, monitored with Prometheus/Grafana, and ran A/B tests for data-driven decisions.
Build and maintain a payment and wallet system serving 51+ million Persian-speaking users globally with high reliability, security, and performance requirements.
Developed payment processing services with careful error handling and retries, implemented secure wallet management with audit trails, designed idempotent APIs to handle network issues, and built real-time monitoring dashboards for transaction tracking.
User-review moderation ran on a low-accuracy text-mining model, so flagging problematic reviews took Bazaar teams roughly two days per batch. Separately, as LLMs matured, there was no established way to get them into the engineering org's daily workflow.
Drove LLM adoption at Cafe Bazaar earlier than most of the industry: replaced the legacy review-moderation model with an LLM-based pipeline, and rolled out Claude Code across the engineering org as a shared development tool. A practical internal web UI let any team use LLMs without deep AI expertise.
Let's discuss how we can architect scalable systems and build high-performing teams together.
Get in Touch