lightseekorg/tokenspeed

GitHub: lightseekorg/tokenspeed

专为智能体工作负载设计的高性能大模型推理引擎,兼具 TensorRT-LLM 级性能与 vLLM 级易用性。

Stars: 1792 | Forks: 213

TokenSpeed: Tokens at the speed of light

TokenSpeed 是一款专为**智能体工作负载**设计的光速 LLM 推理引擎,具备 TensorRT-LLM 级别的性能和 vLLM 级别的易用性。我们的目标是成为生产级智能体工作负载中性能最强的推理引擎。 核心组件: - **建模层**:采用 local-SPMD 设计以及静态编译器,该编译器可根据模块边界的布局注解生成集合通信,因此用户无需手写并行逻辑。 - **调度器**:C++ 控制平面和 Python 执行平面。请求生命周期、KV cache 所有权以及重叠时序均被编码为有限状态机,并在编译时由类型系统强制执行安全的 KV 资源重用。 - **Kernels**:可插拔的分层 kernel 系统,带有可移植的公共 API 和集中式注册表,其中包含了 Blackwell 平台上针对智能体工作负载最快的 **MLA**(Multi-head Latent Attention)实现之一。 - **入口点**:集成 SMG 的 AsyncLLM,用于实现低开销的 CPU 端请求处理。 ## 新闻 - [2026/07] Day 0 上的 [Kimi K3](https://huggingface.co/moonshotai/Kimi-K3#5-deployment):基于 TokenSpeed 在主流平台上实现前沿模型支持。[[博客](https://lightseek.org/blog/tokenspeed-kimi-k3.html)] - [2026/07] Day 0 上的 [TML Inkling](https://thinkingmachines.ai/news/introducing-inkling/):在 NVIDIA 和 [AMD](https://huggingface.co/lightseekorg/Inkling-MXFP4) 上借助 [TokenSpeed](https://thinkingmachines.ai/news/introducing-inkling/#inkling-availability) 进行 FP4 推理。[[博客](https://lightseek.org/blog/tokenspeed-inkling.html)] - [2026/06] 深入探讨 TokenSpeed-Kernel 的设计与优化。[[博客](https://pytorch.org/blog/lightseek-tokenspeed-kernel/)] - [2026/05] 🚀 TokenSpeed 在 Qwen3.5-397B-A17B 上的智能体工作负载吞吐量达到 580 TPS。[[博客](https://pytorch.org/blog/up-to-580tps-new-speed-record-of-qwen3-5-397b-a17b-on-gpu-for-agentic-workloads-with-tokenspeed/)] - [2026/05] TokenSpeed 发布 —— 一款专为智能体工作负载设计的光速 LLM 推理引擎。[[博客](https://lightseek.org/blog/lightseek-tokenspeed.html)] ## 博客与演讲 如需了解来自 LightSeek Foundation 的技术博客、大会演讲及工程文章,请访问 [LightSeek 博客](https://lightseek.org/blog/)。 ## 性能对比 TokenSpeed 与 TensorRT-LLM 在智能体工作负载上的 Pareto 曲线对比 (Kimi K2.5, B200) ## 文档 从这里开始: - [文档索引](https://lightseek.org/tokenspeed/) - [快速入门](https://lightseek.org/tokenspeed/guides/getting-started) - [启动服务器](https://lightseek.org/tokenspeed/guides/launching) - [模型配置方案](https://lightseek.org/tokenspeed/recipes/models) - [服务器参数](https://lightseek.org/tokenspeed/configuration/server) - [兼容参数](https://lightseek.org/tokenspeed/configuration/compatible-parameters) - [并行策略](https://lightseek.org/tokenspeed/serving/parallelism)
标签:Vectored Exception Handling, 凭据扫描, 逆向工具