Principal AI Agent Engineer
Profile
10 years across large-scale platforms and AI engineering at Alibaba and SenseTime, now leading AI Agent platform architecture, productization, production delivery, and team AI engineering effectiveness.
- Service scale
- 2k+ tenants / 230k+ users
- Peak traffic
- 60k / day
- Team scope
- 20-person product-engineering / 4-person AI Agent squad
- Delivery efficiency
- 40 repos / ~10×
Experience
Owned product architecture and production operations for a cross-industry WhatsApp AI Agent SaaS serving 2k+ tenants, 230k+ users, and 60k peak requests/day across business AI Agents, RAG/Memory, tools and Skills, quality, cost, and team delivery.
Evolved WhatsApp customer service from rule flows to tenant-configurable LLM-driven business AI Agents across intent routing, SOPs, knowledge, tools, callbacks, and handoff; delivered 70%+ auto-resolution, about 7s end-to-end latency at 7k+ tokens per conversation, and 60% lower token use.
Built a unified knowledge and memory layer for multi-tenant business AI Agents, injecting trusted knowledge, conversation facts, user preferences, and task state on demand.
- RAG
- Unified ingestion, chunking, embeddings, retrieval, and reranking for documents, QA, websites, attachments, and product data; added model-driven Agentic RAG for retrieval planning, query rewriting, source selection, and iterative evidence gathering.
- Memory
- Built tenant-isolated session and user Memory with fact extraction, confidence filtering, write/update, and retrieval injection for persistent conclusions, preferences, and task state.
Led model-driven AI Agent 1.0 from 0 to 1 in a SaaS business, evolving a Dify workflow into a controllable multi-tenant AI Agent platform for marketing, ecommerce shopping assistance, lead identification, data collection, qualification, business queries, and human handoff, with repeatable delivery across industries.
- Configuration and controlled reasoning
- Built a configuration-driven Agent factory with tenant- and Agent-level customization across Instructions, Actions, knowledge bases, Skills, MCP, language style, and human handoff; used Guidelines, Journeys, and ARQs to govern rules, multi-turn workflows, and critical decisions.
- Engineering controls and runtime assurance
- Used an Agent Harness to standardize engineering rules, tool boundaries, recovery, and release validation, backed by an observability console for conversation flows, tool calls, quality, latency, and cost.
- Quality operations
- Established pre-release evaluation, production replay, issue attribution, and cost analysis, comparing task success, tool accuracy, latency, and token cost by tenant, scenario, and version to support continuous optimization and 70%+ auto-resolution.
Led an internal AI engineering delivery harness and CLI platform across 40 repositories, using shared context to connect development, testing, release, deployment, verification, and operations while consolidating a long multi-role workflow into a developer-led AI delivery loop.
- Engineering and testing convergence
- Unified requirements, implementation, unit and integration tests, case reproduction, and pre-acceptance checks through a shared CLI, reusable Skills, and engineering checks; engineers owned and automated most test execution while QA focused on high-risk scenarios and quality gates.
- End-to-end AI delivery loop
- Connected build, release, deployment, environment checks, regression verification, production observability, and incident diagnosis through AI-driven commands and Skills that execute and write results back into shared context.
- Collaboration efficiency
- Shared repository rules, service dependencies, environment configuration, and Spec/EPIC context through Skills, MCP, and CLI, reducing handoffs, verbal synchronization, and wait states; standardized delivery and production issue resolution reached roughly 10x the prior collaboration model's efficiency.
Built a multilingual unknown-question clustering and knowledge-gap learning platform that turned unresolved questions into actionable themes and candidate knowledge across discovery, synthesis, tenant adoption, and write-back.
- Algorithm pipeline
- Unified multilingual cleaning, embeddings, UMAP/HDBSCAN, BERTopic/c-TF-IDF, and LLM clustering with SentenceTransformer and OpenAI-compatible interfaces.
- Engineering platform
- Built an async Python, FastAPI, Pydantic, and MongoDB platform with concurrency controls, dynamic configuration, webhook retries, experiment tracking, and quality metrics.
- Knowledge evolution and outcomes
- Evolved HDBSCAN into LLM clustering to identify knowledge gaps and learn candidate answers from high-quality human conversations; about 50% of clusters entered production optimization, high-frequency coverage reached 70%, and auto-resolution improved 30%.
Built and delivered independent AI data and model-serving products, owning product definition, system architecture, full-stack implementation, containerized deployment, and production iteration.
Designed a controlled conversational Agent framework from 0 to 1 for enterprise consultation and complex tasks, separating business rules, workflow state, tool execution, and response generation into governable layers for reliable long-running conversations.
- Rules and workflow model
- Modeled local business rules as condition-action Guidelines selected by context, while Journeys carried branching and backtracking multi-turn flows between free-form conversation and deterministic SOPs.
- Attentive reasoning
- Applied Attentive Reasoning Queries (ARQs) as domain reasoning blueprints that reinstate constraints around intent, rule applicability, tool choice, parameter provenance, and response compliance.
- Production execution framework
- Built a configuration-driven Agent factory, session runtime, RAG and tool integration, async execution, evaluation, and observability, bringing rules, tools, and model versions into release governance and replay.
Built a production multimodal data and model platform spanning annotation, model artifacts, auto-labeling, quality validation, and pluggable Nuclio/Traefik inference for audio, video, and image workflows.
Owned backend engineering and microservice architecture for AI image generation, model training, and platform products, contributing to 0-to-1 generative AI SaaS delivery.
Owned engineering integration and service architecture for proprietary image models, LoRA training, and open-source models, covering parameter adaptation, prompt processing, async queues, model storage, post-processing, trace logs, and monitoring; defined RPC, load balancing, and throughput optimization for stable Stable Diffusion/PyTorch production serving.
Worked across Alibaba Local Services commercial promotion and Taobao new-retail engineering on high-traffic cross-platform systems spanning ad delivery and attribution, live interaction, micro-frontends, SSR, and performance.
Moved Taobao Live interaction from isolated Weex components to H5 micro-frontends with unified routing and communication; combined SSR and tiered resources to keep FCP below 1.5s under high traffic and weak networks.
Owned financial reporting, Java web services, and frontend migration from AngularJS to Vue with a shared Web UI component library.
Built Spring Boot/MyBatis financial operations services, migrated AngularJS to Vue, and established a shared Web UI library.
Skills
Agent Harness / Loop Engineering / Agentic RAG & Memory / Multi-agent Orchestration / Tool-use & MCP / Skills / Evaluation & Observability / Context Engineering / Model Serving, Routing & LLMOps
Embeddings / Reranking / UMAP / HDBSCAN / BERTopic & c-TF-IDF / LLM Clustering / Knowledge-gap Learning
Python / FastAPI / TypeScript / Rust / Java / PostgreSQL / Redis / Milvus / Docker / Kubernetes / Microservices / CI/CD / Alibaba Cloud / Google Cloud