From AI Demo to Production: Engineering Practices for Agent and AI Native Applications
This event explores the engineering challenges of turning AI demos into sustainable products, featuring discussions on persistent agent memory, robust backend workflows, and efficient model routing.

Event Background
Building a working AI demo is no longer difficult; the real challenge is turning it into a product that can run, iterate, and be delivered continuously.
Persistent Memory: Ensuring project context, historical decisions, and task progress persist across sessions, tasks, and agents. Controllable Backend: Enabling stable generation, maintenance, and handover of databases, APIs, workflows, and role-based permissions. Efficient Invocation: Balancing model performance, latency, success rate, and cost in multi-model and high-frequency invocation scenarios.
This event invites frontline practitioners from MemTensor, Zion, PPIO, and FastGPT to discuss cross-agent long-term memory, Vibe Coding backend development, and intelligent model gateways, sharing architectural designs and engineering practices for transitioning AI applications from tech demos to real products.
Talk Details
Make All AI Remember the Same You: Memory Architecture for Memmy & MemOS
Zong Yue (AI Full-Stack Developer, Memmy Tech PM)
AI full-stack developer and Memmy Tech PM, responsible for both requirements and technical solution design, and personally involved in the code implementation of Agent Runtime, cross-agent memory integration, local services, and cloud business. Continuously exploring the engineering path of Agent Memory from tech demo to real product.
As agents like Cursor, Claude Code, and Codex gradually enter development workflows, project context, historical decisions, and task progress begin to scatter across different tools. How to break down the memory silos between tools so that context can continuously persist across sessions, tasks, and agents has become a critical problem that must be solved for agents to move from single invocation to long-term operation.
This talk will start from MemTensor’s proprietary Agent Memory algorithm, introducing how MemOS provides memory writing, organization, retrieval, update, and governance capabilities for AI applications and agent systems. It will also cover how Memmy leverages MemOS’s memory capabilities to precipitate multi-agent history into a sustainable personal long-term memory. The content will include Memmy’s product design, system architecture, history import, and task continuation, dissecting the complete chain of cross-agent long-term memory from underlying capabilities to product landing through real demos.
Smarter Tokens, Cheaper Intelligence: PPIO Intelligent Model Gateway
Chen Jiaqi
As large model applications enter the agentification stage, model invocation systems face new engineering challenges: How to dynamically route across multiple models? How to determine which tasks require strong models and which can be handled by lightweight models? How to maintain low latency, high success rates, and controllable costs under high concurrency and high-frequency invocation? These questions have exceeded the capabilities of traditional API gateways.
This session will explore from an engineering implementation perspective, introducing PPIO intelligent model gateway’s design concepts in MoM (Mixture of Models) inference, intelligent routing, and context compression. The talk will combine typical scenarios such as in-depth research, enterprise knowledge Q&A, content generation, and code assistance, to demonstrate how intelligent model gateways abstract complex model selection, invocation optimization, and quality control capabilities into a unified service. This allows developers to focus on Agent business logic and provides Agents with the foundation for long-term, low-cost, and scalable operation.
Rescuing Vibe Coding Developers Stuck on Backend
Qin Mao (Tim), Zion
More people without engineering backgrounds are using AI programming tools to turn their ideas into working demos—describe a page or interaction, and AI generates the code to see it run.
But at a certain stage, these people almost all hit the same wall: pages are easy to generate, but backend business logic is very difficult to stably hand over to AI for generation, maintenance, and transition.
Zion provides a development plugin and platform Copilot for this, allowing users to generate clear and understandable databases, API configurations, complex business logic workflows, AI agents, and role permission management through natural language descriptions, and automatically deploy them to the cloud, making backend development controllable and sustainably iterative.
Making Agents Truly Work: From Models and Memory to Deliverable Applications
FastGPT Solution Lead
FastGPT Solution Lead, focusing long-term on the application landing and solution design of enterprise-level AI Agents, dedicated to transforming capabilities such as large models, knowledge bases, workflows, and tool calls into deployable, deliverable, and sustainably iterative AI applications.
Models give agents the ability to think, memory gives them context, and tools give them the ability to act. However, to make these capabilities truly serve users, an application layer is needed to organize, orchestrate, and deploy them.
This talk explores how FastGPT connects models, knowledge bases, contexts, workflows, and tool calling to help developers progressively transform a working demo into a debuggable, deployable, and deliverable agent application.
Programme
Registration & Networking
Check-in and open networking for developers.
Keynote: Make All AI Remember the Same You
Memory architecture and engineering practices for Memmy and MemOS.
Keynote: Smarter Tokens, Cheaper Intelligence
PPIO's intelligent model gateway practices.
Keynote: Rescuing Vibe Coding Developers Stuck on Backend
Making backend development controllable and sustainable.
Keynote: Making Agents Truly Work
From models and memory to deliverable applications.
