The Evolution of Next-Generation AI Inference Infrastructure (Tentative) — Jooyoung Kim | AI Summit Seoul 2026

The Evolution of Next-Generation AI Inference Infrastructure (Tentative)

Session Overview

As large language models become more widely adopted, the focus of the AI industry is shifting beyond model training toward how quickly and efficiently those models can be executed in real-world services. With the rapid growth of generative AI and AI agents, power consumption, infrastructure cost, processing speed, and scalability during inference are becoming critical factors that directly affect business competitiveness.

This session explores the evolution of next-generation AI inference infrastructure, with a focus on the LPU (LLM Processing Unit), a processor designed specifically for LLM inference. It will examine the efficiency and cost limitations of conventional GPU-centered computing architectures and introduce approaches to improving performance and energy efficiency by jointly optimizing AI semiconductors, memory, and system architecture.

The session will also present HyperAccel’s high-efficiency inference accelerator, built on Samsung’s 4-nanometer process and LPDDR5X memory, as a case study in achieving the performance, cost efficiency, and scalability required for AI data centers and enterprise environments. It will further consider the role that dedicated inference semiconductors will play as AI models and services continue to advance.

Key Takeaways

  • The growing importance of inference infrastructure as generative AI and AI agents expand
  • The power, cost, and scalability limitations of GPU-centered computing architectures
  • How LPU architecture improves performance and energy efficiency for LLM inference
  • Strategies for jointly optimizing semiconductors, memory, and system architecture
  • How dedicated inference processors will reshape next-generation AI data centers and enterprise infrastructure

Speaker

Jooyoung Kim
Jooyoung Kim
CEO
HyperAccel
AI Semiconductor LLM Inference AI Infrastructure

Jooyoung Kim founded HyperAccel in 2023 and is leading the development of the HyperAccel LPU (LLM Processing Unit), an AI semiconductor optimized for large language model inference. Through a high-efficiency inference accelerator built on Samsung’s 4-nanometer process and LPDDR5X memory, he is focused on building next-generation AI infrastructure with significantly improved power and cost efficiency.

He has more than 20 years of research and industry experience across AI semiconductor design, system architecture, computing platforms, and product commercialization. Drawing on this background, he leads the research, development, and commercialization of next-generation AI computing technologies.

Before founding HyperAccel, he spent nine years at Microsoft’s U.S. headquarters, where he led research and development in AI accelerators and computing infrastructure. Since 2019, he has also served as a professor in the School of Electrical Engineering at KAIST.

Register Now
Program and speaker information may be updated prior to the event.