Premium Inference: Why Speed, Cost, and Tokens Are the New Battleground of Enterprise AI
Inference is becoming the real bottleneck of production AI. SambaNova's Reggie Lu brings the blueprint for the next generation of AI infrastructure to AISE 2026.
Training a great model gets the headlines. Serving it — fast, securely, and at a cost that makes business sense — is where enterprise AI lives or dies. At AI Summit Seoul & Expo 2026, Reggie Lu, Principal Customer Engineer APAC at SambaNova, will introduce premium inference as the next frontier of AI infrastructure.
The session
In "Delivering the Blueprint for Premium Inference: From Faster Tokens to Disaggregated AI Infrastructure," Lu makes the case that as enterprises adopt coding assistants, multimodal AI, and agentic workflows, success depends not only on model quality but on delivering fast, reliable, secure, and cost-efficient intelligence at scale. Latency, throughput, concurrency, memory efficiency, and cost per useful token now matter more than ever.
The session also explores disaggregated inference — separating prefill, decode, and orchestration across the right infrastructure layers — as a new blueprint for scaling enterprise AI. It's designed for anyone who wants to understand how the next generation of AI infrastructure will power real-world, production-grade systems.
Why it matters now
Agentic AI has changed the economics of inference. When autonomous agents run continuously — calling models thousands of times per task rather than once per prompt — inefficient serving quietly becomes one of the largest line items in an AI budget. Across APAC, where sovereign AI infrastructure is being built at unprecedented pace, the question of how to serve intelligence efficiently is now a board-level concern. Lu's session arrives exactly as the industry's focus shifts from "can we build it?" to "can we afford to run it?"
About the speaker
Reggie Lu is an AI infrastructure and GenAI specialist at SambaNova, based in Tokyo. He helps enterprises design, deploy, and scale production-grade AI systems, with deep expertise in LLM inference, enterprise RAG, agentic AI applications, and secure on-premises AI infrastructure. His background spans end-to-end AI productization — from architecture design and model serving to benchmarking, optimized deployment, and production operations — built on years as an ML and AI infrastructure software engineer.
Key Takeaways
- The new metrics that matter — latency, throughput, concurrency, memory efficiency, and cost per useful token as the KPIs of enterprise AI.
- Disaggregated inference, explained — how separating prefill, decode, and orchestration unlocks performance and cost efficiency.
- From experimentation to production — a practical view of what infrastructure the agentic era actually requires.
Session at a Glance
Early Bird pricing ends June 30 · August 19–20, 2026 · COEX, Seoul
Get Your TicketsSubscribe to our newsletter
and be the first to get updates on AI Summit Seoul & Expo 2026!
Stay informed with the latest global AI news and related events.

