Skip to main content

KTransformers Tutorial: Run GLM-5.3-Flash Locally

Run a massive MoE model locally by combining GPU VRAM, system RAM, SGLang, and KTransformers for heterogeneous CPU-GPU inference.
Oct 6, 2026  · 11 min read

Explore with AI

ChatGPTClaudePerplexity

KTransformers is an open-source inference framework that lets the CPU and GPU actively execute different experts during inference, so you can run Mixture-of-Experts (MoE) models that are far larger than your GPU memory. Frameworks like vLLM can also offload weights to CPU memory, but KTransformers is built specifically around the sparse structure of MoE models.

In this tutorial, we'll use KTransformers and SGLang to run GLM-5.3-Flash, a 320B-parameter model whose weights don't fit into 192 GB of VRAM. We'll monitor CPU and GPU memory usage, experiment with expert placement, test the OpenAI-compatible API, and connect the model to Pi as a local coding agent.

The main idea is simple: instead of treating CPU memory as overflow storage, KTransformers uses both CPU compute and GPU compute during inference.

In a Nutshell

  • KTransformers runs large MoE models across GPU VRAM and system RAM, with the CPU computing the experts that stay in RAM.
  • GLM-5.3-Flash's native FP8 weights take up about 306 GiB, so the official guide recommends at least 350 GB of available system memory.
  • We ran the full model on 2× RTX PRO 6000 GPUs (192 GB of VRAM in total) with a 32K context window, at around 11 tokens per second.
  • The server exposes an OpenAI-compatible API, so coding agents like Pi can use the model directly.

What Is KTransformers?

KTransformers is an open-source inference framework for running very large language models using a combination of GPU VRAM and CPU RAM. Normally, serving a large model requires loading most of its weights into GPU memory, which quickly becomes expensive for a model the size of GLM-5.3-Flash.

KTransformers takes a different approach: it keeps many MoE expert weights in system memory and reserves GPU memory for the parts of inference that benefit most from GPU acceleration.

This works well for MoE models because not every expert is used for every token. GLM-5.3-Flash, for example, has 288 routed experts, but its router selects only 8 of them (plus 1 shared expert) for each token. KTransformers can therefore distribute expert computation across the CPU and GPU:

KTransformers workflow diagram showing MoE experts split between CPU and GPU

How KT-Kernel and SGLang work together

The current KTransformers stack integrates KT-Kernel with SGLang for heterogeneous CPU-GPU inference. Each component handles a different job:

  • SGLang provides the serving runtime: API requests, batching, request scheduling, KV-cache management, and GPU parallelism.
  • KT-Kernel replaces the standard MoE execution path with CPU-GPU-aware expert execution. Selected experts run on the GPU, while the rest stay in CPU memory and are computed on the CPU.

KTransformers also supports changing expert placement based on workload patterns, as described in its expert scheduling tutorial.

In other words, KTransformers treats CPU memory and GPU memory as a shared inference system rather than requiring the entire model to fit inside GPU VRAM. That is what lets very large MoE models run on hardware with far less GPU memory than they would normally need.

What Is GLM-5.3-Flash?

GLM-5.3-Flash is Z.ai's open-weight, natively multimodal MoE model, released under the MIT license in August 2026. Despite the "Flash" name, it's a large model: 320B total parameters, with about 18B active per token.

These are the specs that matter for local inference:

  • Experts: 288 routed experts with top-8 routing, plus 1 shared expert
  • Weights: about 306 GiB for the official FP8 checkpoint (zai-org/GLM-5.3-Flash)
  • Context window: up to 1M tokens
  • Inputs: text, images, and video, with support for reasoning and tool calling

KTransformers reads the official FP8 weights directly, so there's no conversion or extra quantization step. For benchmarks and a full model overview, see our GLM-5.3-Flash guide.

GLM-5.3-Flash Hardware Requirements

For GLM-5.3-Flash, the hardware question is mostly about system RAM. The official KTransformers GLM-5.3-Flash tutorial recommends reserving at least 350 GB of available system memory.

For this tutorial, we use a RunPod instance with approximately:

GPU:         2× RTX PRO 6000
VRAM:        96 GB each
Total VRAM:  192 GB

System RAM:  350 GB+
Storage:     500 GB+
Python:      3.11

Launching a RunPod pod with 2× RTX PRO 6000 GPUs

GLM-5.3-Flash's official FP8 checkpoint is approximately 306 GiB (about 329 GB), while our two GPUs provide 192 GB of VRAM in total. The full model therefore cannot simply be loaded into GPU memory.

Instead, KTransformers keeps a large portion of the MoE weights in system RAM and moves the most useful computation to the GPUs. The 350 GB recommendation leaves enough room for the model weights plus runtime overhead.

More GPU memory does not remove the need for RAM in this setup. CPU memory is an intentional part of KTransformers' heterogeneous inference design: expert weights remain in RAM while the GPU handles the parts of the model that benefit most from acceleration.

The current GLM-5.3-Flash implementation also has specific CPU and GPU requirements:

  • GPU: NVIDIA SM89 or SM120 architectures, which covers the RTX 40 series, RTX 50 series, and Blackwell workstation cards such as the RTX PRO 6000.
  • CPU: AVX-512 support, which the FP8 CPU expert kernel relies on.

Can you run GLM-5.3-Flash on a single GPU?

Yes, provided you have enough system RAM and a supported CPU. The official tutorial includes a single-GPU configuration that sets --kt-num-gpu-experts 0, so the MoE experts are handled on the CPU side.

We use two RTX PRO 6000 GPUs here, but that is not a strict minimum requirement. The second GPU gives us more VRAM and extra headroom while experimenting with a relatively new KTransformers implementation, rather than tuning the setup around the smallest hardware configuration that can possibly run the model.

Step 1: Install KTransformers With SGLang

Create a clean Python 3.11 environment and install KTransformers with SGLang support:

python3.11 -m venv /workspace/kt
source /workspace/kt/bin/activate

pip install --upgrade pip
pip install "ktransformers[sglang]"

Verify that KTransformers, KT-Kernel, SGLang, and CUDA are detected correctly:

kt version

You should see output similar to:

KTransformers CLI v0.7.0.post4

Python      3.11.13
Platform    Linux 6.8.0-136-generic
CUDA        13.0

Packages:
kt-kernel   0.7.0.post4
sglang-kt   0.7.0.post4

This confirms that the KTransformers runtime and its SGLang backend are installed and ready to use.

Step 2: Download GLM-5.3-Flash From Hugging Face

Before starting the server, download the official GLM-5.3-Flash checkpoint from Hugging Face:

hf download zai-org/GLM-5.3-Flash \
  --local-dir /workspace/GLM-5.3-Flash

Downloading the zai-org/GLM-5.3-Flash model from Hugging Face

The checkpoint is around 306 GiB, so the download can take a while depending on your bandwidth.

Then point KTransformers to the local model path:

export MODEL_PATH=/workspace/GLM-5.3-Flash

Step 3: Launch the GLM-5.3-Flash Server With SGLang

Now launch GLM-5.3-Flash with two-way tensor parallelism, using both RTX PRO 6000 GPUs. The model supports up to 1M tokens of context, and the official examples use a validated 501,025-token configuration. We start with a 32K context window instead, to keep memory usage predictable while testing the setup.

CUDA_VISIBLE_DEVICES=0,1 \
python -m sglang.launch_server \
  --model-path "$MODEL_PATH" \
  --kt-weight-path "$MODEL_PATH" \
  --served-model-name GLM-5.3-flash \
  --host 0.0.0.0 \
  --port 30000 \
  --tp-size 2 \
  --context-length 32768 \
  --max-total-tokens 32768 \
  --mem-fraction-static 0.85 \
  --chunked-prefill-size 2048 \
  --kt-method FP8 \
  --kt-cpuinfer 64 \
  --kt-threadpool-count 2 \
  --kt-num-gpu-experts 14 \
  --kt-gpu-prefill-token-threshold 2048 \
  --kt-expert-placement-strategy uniform \
  --cuda-graph-bs 1 2 4 \
  --enable-p2p-check \
  --tool-call-parser glm47 \
  --reasoning-parser glm45

Running the zai-org/GLM-5.3-Flash model with SGLang and KTransformers

This configuration exposes the model through an OpenAI-compatible SGLang server on port 30000. The two GPUs are used with --tp-size 2, while KTransformers keeps part of the MoE workload on the CPU and places selected experts on the GPUs.

The settings here are intentionally conservative for the first run: 32K context, 85% static GPU memory usage, 14 GPU experts, and 64 CPU inference threads. Once the server is stable, you can experiment with a larger context window, more GPU experts, or different memory settings to improve throughput.

Key KTransformers launch flags explained

Most of the flags above are standard SGLang options. These are the ones that control how KTransformers splits work between the CPU and GPU:

Flag Value What it does
--kt-method FP8 Sets the precision of the expert weights, matching GLM-5.3-Flash's native FP8 checkpoint.
--kt-cpuinfer 64 Number of CPU threads used for expert computation.
--kt-threadpool-count 2 Number of CPU thread pools, usually matched to the number of NUMA nodes.
--kt-num-gpu-experts 14 Number of experts per MoE layer placed on the GPU.
--kt-expert-placement-strategy uniform How GPU experts are chosen. Other options include frequency, front-loading, and random.
--kt-gpu-prefill-token-threshold 2048 Prompt length above which prefill switches to the GPU-side layerwise path.

Step 4: Test CPU-GPU Offloading and Expert Placement

Now that the server is running, we can verify how KTransformers is using GPU VRAM and system RAM, then change the number of GPU-resident experts to see how resource usage and performance shift.

On RunPod, free -h can be misleading because a container may see the host machine's total RAM rather than only the memory available to the pod. It is better to monitor GPU memory and container memory separately.

Monitor GPU VRAM usage

Open a new terminal and monitor GPU usage:

watch -n 1 nvidia-smi

Monitoring GPU VRAM usage while running GLM-5.3-Flash with KTransformers

With our current configuration, the fully loaded model uses roughly 48 GB per GPU, leaving a large amount of VRAM unused. That suggests there is room to place more experts on the GPUs, or to test a single-GPU configuration with sufficient system RAM.

Monitor container RAM on RunPod

For container RAM, read the cgroup memory counters directly:

watch -n 1 'echo -n "Used: "; awk "{printf \"%.1f GiB\n\", \$1/1024/1024/1024}" /sys/fs/cgroup/memory.current; echo -n "Limit: "; awk "{printf \"%.1f GiB\n\", \$1/1024/1024/1024}" /sys/fs/cgroup/memory.max'

Monitoring container system memory usage on RunPod

You should see that a large portion of the available system RAM is occupied by model weights and CPU-side experts. This is expected: KTransformers deliberately keeps many MoE experts in RAM instead of requiring them all to live in VRAM.

Tune --kt-num-gpu-experts

Next, restart the server with different values of --kt-num-gpu-experts. For example, compare:

0
10
20

--kt-num-gpu-experts controls how many experts per MoE layer are placed on the GPU. With 0, expert computation stays on the CPU side; increasing the value moves more experts into GPU memory.

For each configuration, compare GPU VRAM usage, container RAM usage, tokens per second, and time to first token. In general, more GPU experts consume more VRAM but reduce CPU-side expert execution, which can improve inference performance when enough VRAM is available.

One caveat from the official tutorial: when Layerwise Prefill is enabled for GLM-5.3-Flash, the current implementation normalizes the resident GPU expert count to zero. If VRAM usage barely changes between runs, this is the likely reason.

This experiment shows the key advantage of KTransformers: CPU RAM and GPU VRAM become tunable parts of the same inference system, so you can trade memory placement against speed instead of requiring the entire MoE model to fit on the GPUs.

Step 5: Test the OpenAI-Compatible API

With the server running, we can now confirm that the model is available and send a real request through SGLang's OpenAI-compatible API.

First, check that the model is registered:

curl http://localhost:30000/v1/models

Next, send a test prompt:

curl http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3-flash",
    "messages": [
      {
        "role": "user",
        "content": "Create a FastAPI application with a health endpoint."
      }
    ],
    "max_tokens": 500
  }'

Chat completion response generated by GLM-5.3-Flash through the KTransformers OpenAI-compatible API

If the setup is working correctly, the server returns a normal chat completion response containing generated code and usage statistics.

Step 6: Use GLM-5.3-Flash as a Local Coding Agent With Pi

Pi is a lightweight coding agent that can use any OpenAI-compatible model as its backend, which lets GLM-5.3-Flash work directly on coding tasks instead of only answering prompts.

Install Pi

Install Pi with its install script:

curl -fsSL https://pi.dev/install.sh | sh

Installing the Pi coding agent

Then add Pi to your PATH:

echo 'export PATH="/root/.local/share/pi-node/current/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc

Point Pi at the KTransformers server

Create a model configuration that points Pi to the local KTransformers server:

mkdir -p ~/.pi/agent && cat > ~/.pi/agent/models.json <<'EOF'
{
  "providers": {
    "ktransformers": {
      "baseUrl": "http://localhost:30000/v1",
      "api": "openai-completions",
      "apiKey": "local",
      "models": [
        {
          "id": "GLM-5.3-flash",
          "name": "GLM-5.3-Flash",
          "reasoning": true,
          "input": ["text"],
          "contextWindow": 32768,
          "maxTokens": 8192,
          "cost": {
            "input": 0,
            "output": 0,
            "cacheRead": 0,
            "cacheWrite": 0
          }
        }
      ]
    }
  }
}
EOF

Start Pi:

pi

Then open the model selector:

/model

Selecting the KTransformers-served GLM-5.3-Flash model in Pi

Run a coding task with GLM-5.3-Flash

Choose GLM-5.3-Flash and try a real coding task:

Build a FastAPI service with /health and /users endpoints.
Add pytest tests and run them.

GLM-5.3-Flash working on a FastAPI coding task in the Pi coding agent

Within a few seconds, Pi should begin creating files, writing the API, running tests, and fixing problems as it works through the task.

Monitoring the SGLang and KTransformers inference server logs

You can also watch the first terminal, where the SGLang server is running. In our test, generation speed was around 11 tokens per second. That is reasonable for a model of this size running with substantial CPU offloading, and the setup can be tuned further by moving more experts onto the GPUs.

Pi coding agent summary after GLM-5.3-Flash completed the FastAPI task

Within a few minutes, the model had created the endpoints, written and executed the tests, performed a smoke test, and produced a short summary explaining how to run the project.

The interesting part is that the complete model is running locally even though its weights are much larger than the available GPU VRAM. Pi handles the coding-agent loop, while SGLang and KTransformers handle the actual model inference.

KTransformers vs vLLM vs llama.cpp

KTransformers is not the only way to run a model that is larger than your VRAM. vLLM and llama.cpp both support CPU offloading, but they split the work differently:

Framework How it uses CPU memory Where expert computation runs Best fit
vLLM Offloads part of the weights to CPU RAM (--cpu-offload-gb) and transfers them to the GPU when needed GPU High-throughput serving when the model mostly fits in VRAM
llama.cpp Splits layers between CPU and GPU, and can keep MoE expert tensors in RAM (--n-cpu-moe) CPU and GPU Quantized GGUF models on consumer hardware
KTransformers + SGLang Keeps most experts in RAM and places a set number of experts per layer on the GPU CPU and GPU, with optimized AVX-512 expert kernels Native-precision MoE models on machines with hundreds of GB of RAM

SGLang and KTransformers are not competing in this setup. SGLang handles the serving side, while KTransformers handles the heterogeneous CPU-GPU MoE execution.

Final Thoughts

What I liked most about this setup is that KTransformers does something a little different from the usual inference stack. Instead of thinking only in terms of model layers, it places individual experts on the GPU while keeping others in system RAM, and the CPU actually takes part in expert computation. The CPU is not just acting as overflow storage.

In this tutorial, we ran the full GLM-5.3-Flash model locally on two RTX PRO 6000 GPUs, even though its weights are much larger than the available VRAM.

It is not the fastest setup. I was getting around 11 tokens per second, and there is still a lot of room to tune the number and placement of GPU experts. You could also experiment with a single GPU if you have enough RAM; I used two here simply to give the setup more headroom.

For me, that is the main takeaway from this guide: KTransformers is not special because it invented CPU offloading. It is special because it makes CPU RAM, CPU compute, and GPU compute work together around the sparse structure of MoE models.

KTransformers and GLM-5.3-Flash FAQs

How much RAM do you need to run GLM-5.3-Flash with KTransformers?

The official KTransformers tutorial recommends at least 350 GB of available system memory. The native FP8 weights take up about 306 GiB, and the rest covers runtime overhead.

Can KTransformers run GLM-5.3-Flash on a single GPU?

Yes. The official tutorial includes a single-GPU configuration with --kt-num-gpu-experts 0, which keeps expert computation on the CPU. You still need enough system RAM and a CPU with AVX-512 support.

Which GPUs and CPUs does KTransformers support for GLM-5.3-Flash?

The current implementation supports NVIDIA SM89 and SM120 GPUs, which includes the RTX 40 series, RTX 50 series, and the RTX PRO 6000. On the CPU side, the FP8 expert kernel requires AVX-512.

How fast is GLM-5.3-Flash with KTransformers?

In our test on 2× RTX PRO 6000 GPUs with a 32K context window and 14 GPU experts per layer, generation speed was around 11 tokens per second. Speed depends mainly on how many experts sit on the GPU, your CPU, and your memory bandwidth.

Which other models does KTransformers support?

KTransformers supports a range of large MoE models, including GLM-5, GLM-5.2, Kimi K2.5, MiniMax-M2.5, and Qwen3-235B-A22B. Check the KTransformers GitHub repository for the current list and model-specific tutorials.


Abid Ali Awan's photo
Author
Abid Ali Awan
LinkedIn
Twitter

As a certified data scientist, I am passionate about leveraging cutting-edge technology to create innovative machine learning applications. With a strong background in speech recognition, data analysis and reporting, MLOps, conversational AI, and NLP, I have honed my skills in developing intelligent systems that can make a real impact. In addition to my technical expertise, I am also a skilled communicator with a talent for distilling complex concepts into clear and concise language. As a result, I have become a sought-after blogger on data science, sharing my insights and experiences with a growing community of fellow data professionals. Currently, I am focusing on content creation and editing, working with large language models to develop powerful and engaging content that can help businesses and individuals alike make the most of their data.

Topics
Artificial Intelligence
Large Language Models

Top DataCamp Courses

Course

Transformer Models with PyTorch

2 hr
9.2K
What makes LLMs tick? Discover how transformers revolutionized text modeling and kickstarted the generative AI boom.
See DetailsRight Arrow
Start Course
See MoreRight Arrow
Related

Tutorial

How to Run GLM 4.7 Flash Locally

Learn how to run GLM-4.7-Flash on an RTX 3090 for fast local inference and integrating with OpenCode to build a fully local automated AI coding agent.
Abid Ali Awan's photo

Abid Ali Awan

11 min

Tutorial

How to Run GLM-4.7 Locally with llama.cpp: A High-Performance Guide

Setting up llama.cpp to run the GLM-4.7 model on a single NVIDIA H100 80GB GPU, achieving up to 20 tokens per second using GPU offloading, Flash Attention, optimized context size, efficient batching, and tuned CPU threading.
Abid Ali Awan's photo

Abid Ali Awan

10 min

Tutorial

Using Claude Code With Ollama Local Models

Run GLM 4.7 Flash locally (RTX 3090) with Claude Code and Ollama in minutes, no cloud, no lock-in, just pure speed and control.
Abid Ali Awan's photo

Abid Ali Awan

8 min

Tutorial

Run GLM-5 Locally For Agentic Coding

Run GLM-5, the best open-weight AI model, on a single GPU with llama.cpp, and connect it to Aider to turn it into a powerful local coding agent.
Abid Ali Awan's photo

Abid Ali Awan

8 min

Tutorial

How to Run GLM 5.1 Locally For Agentic Coding

Learn how to run GLM 5.1 locally on an H100 GPU with llama.cpp, test it, use the WebUI, and integrate Claude Code.
Abid Ali Awan's photo

Abid Ali Awan

10 min

Tutorial

How to Run DeepSeek V4 Flash Locally

Learn how to run the full DeepSeek V4 Flash model on a single GPU using a modified llama.cpp build and a compatible GGUF file in this hands-on tutorial.
Abid Ali Awan's photo

Abid Ali Awan

9 min

See MoreSee More