Track 1 focuses on the Performance Optimization of SGLang Framework Operators on Multiple Chips. With the rapid advancement of Large Language Models (LLMs), efficient inference has become critical for large-scale AI deployment. Operator optimization is a key technical approach to improving model throughput, reducing latency, and enhancing resource utilization. However, various AI chip platforms exhibit significant differences in hardware architecture, memory hierarchies, and programming models, making high-performance cross-platform operator implementation a major challenge.
This track is jointly organized by FlagOS and IEEE, co-organized by SGLang, and supported by sponsors including Iluvatar CoreX, Enflame Technology, Kunlunxin, and more. It is open to AI systems developers worldwide. The competition adopts Triton and its enhanced extension architecture, Triton-TLE (Triton Language Extensions), as the unified language for kernel development. Featuring more than 200 real-world AI inference kernel challenges derived from the SGLang framework, the track offers a continuous challenge focused on multi-chip kernel optimization through standardized correctness verification, multi-chip performance evaluation, and a real-time leaderboard.
The track features three dedicated awards encouraging competition across three dimensions: * Full-Scope Occupied Award: Based on the total number of kernels occupied. * Breakthrough Conquered Award: Based on the number of kernels conquered first. * Extreme Performance Award: Based on single-kernel extreme performance optimization.
These awards provide a comprehensive assessment of teams' technical expertise in low-level kernel development. The event aims to build a repository of reproducible, portable, and deployable high-performance kernel solutions, laying a solid technical foundation for next-generation LLM inference infrastructure.