Skip to content

[Bug] Unrecognized configuration class <class 'sglang.srt.utils.hf_transformers.common._DeepseekV4ConfigAlias'> #33207

Description

@shahizat

Checklist

  • I searched related issues but found no solution.
  • The bug persists in the latest version.
  • Issues without environment info and a minimal reproducible demo are hard to resolve and may receive no feedback.
  • If this is not a bug report but a general question, please start a discussion at https://github.com/sgl-project/sglang/discussions. Otherwise, it will be closed.
  • Please use English. Otherwise, it will be closed.

Describe the bug

Hello, I want to run deepseek-ai/DeepSeek-V4-Flash-0731 on NVIDIA B200 and RTX PRO 6000 Blackwell GPUs, but I'm encountering the following issue.

Unrecognized configuration class <class 'sglang.srt.utils.hf_transformers.common._DeepseekV4ConfigAlias'> 

Reproduction

Command to reproduce:

uv venv .sglang --python 3.12
source .sglagn/bin/activate

uv pip install -U sglang --pre \
  --index-url https://sgl-project.github.io/whl/cu130/ \
  --extra-index-url https://pypi.org/simple \
  --extra-index-url https://download.pytorch.org/whl/cu130 \
  --index-strategy unsafe-best-match

sglang serve \
  --trust-remote-code \
  --model-path deepseek-ai/DeepSeek-V4-Flash-0731 \
  --tp 2 \
  --moe-runner-backend flashinfer_mxfp4 \
  --mem-fraction-static 0.98 \
  --cuda-graph-max-bs-decode 32 \
  --host 0.0.0.0 \
  --port 30000

Environment

Python: 3.12.13 (main, Jul 10 2026, 00:00:00) [GCC 11.5.0 20240719 (Red Hat 11.5.0-14)]
CUDA available: True
GPU 0,1,2,3,4,5,6,7: NVIDIA B200
GPU 0,1,2,3,4,5,6,7 Compute Capability: 10.0
CUDA_HOME: /usr/local/cuda
NVCC: Cuda compilation tools, release 13.3, V13.3.73
CUDA Driver Version: 595.71.05
PyTorch: 2.11.0+cu130
sglang: 0.5.16
sglang-kernel: 0.4.5+cu130
flashinfer_python: 0.6.14
flashinfer_cubin: Module Not Found
flashinfer_jit_cache: Module Not Found
triton: 3.6.0
transformers: 5.12.1
torchao: 0.17.0+cu130
numpy: 2.3.5
aiohttp: 3.14.3
fastapi: 0.141.1
huggingface_hub: 1.26.0
interegular: 0.3.3
modelscope: 1.39.0
orjson: 3.11.9
outlines: 0.1.11
packaging: 26.2
psutil: 7.2.2
pydantic: 2.14.0a1
python-multipart: 0.0.32
pyzmq: 27.1.0
uvicorn: 0.52.0
uvloop: 0.22.1
vllm: Module Not Found
xgrammar: 0.2.1
openai: 2.6.1
tiktoken: 0.13.0
anthropic: 0.120.2
litellm: Module Not Found
torchcodec: 0.11.1+cu130
NVIDIA Topology:
        GPU0    GPU1    GPU2    GPU3    GPU4    GPU5    GPU6    GPU7    NIC0    NIC1    NIC2    NIC3    NIC4    NIC5    NIC6    CPU Affinity    NUMA Affinity   GPU NUMA ID
GPU0     X      NV18    NV18    NV18    NV18    NV18    NV18    NV18    PHB     PHB     PHB     PHB     PHB     PHB     PHB     0-47    0-1             N/A
GPU1    NV18     X      NV18    NV18    NV18    NV18    NV18    NV18    PHB     PHB     PHB     PHB     PHB     PHB     PHB     0-47    0-1             N/A
GPU2    NV18    NV18     X      NV18    NV18    NV18    NV18    NV18    PHB     PHB     PHB     PHB     PHB     PHB     PHB     0-47    0-1             N/A
GPU3    NV18    NV18    NV18     X      NV18    NV18    NV18    NV18    PHB     PHB     PHB     PHB     PHB     PHB     PHB     0-47    0-1             N/A
GPU4    NV18    NV18    NV18    NV18     X      NV18    NV18    NV18    PHB     PHB     PHB     PHB     PHB     PHB     PHB     0-47    0-1             N/A
GPU5    NV18    NV18    NV18    NV18    NV18     X      NV18    NV18    PHB     PHB     PHB     PHB     PHB     PHB     PHB     0-47    0-1             N/A
GPU6    NV18    NV18    NV18    NV18    NV18    NV18     X      NV18    PHB     PHB     PHB     PHB     PHB     PHB     PHB     0-47    0-1             N/A
GPU7    NV18    NV18    NV18    NV18    NV18    NV18    NV18     X      PHB     PHB     PHB     PHB     PHB     PHB     PHB     0-47    0-1             N/A
NIC0    PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB      X      PHB     PHB     PHB     PHB     PHB     PHB
NIC1    PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB      X      PHB     PHB     PHB     PHB     PHB
NIC2    PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB      X      PHB     PHB     PHB     PHB
NIC3    PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB      X      PHB     PHB     PHB
NIC4    PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB      X      PHB     PHB
NIC5    PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB      X      PHB
NIC6    PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB     PHB      X

Legend:

  X    = Self
  SYS  = Connection traversing PCIe as well as the SMP interconnect between NUMA nodes (e.g., QPI/UPI)
  NODE = Connection traversing PCIe as well as the interconnect between PCIe Host Bridges within a NUMA node
  PHB  = Connection traversing PCIe as well as a PCIe Host Bridge (typically the CPU)
  PXB  = Connection traversing multiple PCIe bridges (without traversing the PCIe Host Bridge)
  PIX  = Connection traversing at most a single PCIe bridge
  NV#  = Connection traversing a bonded set of # NVLinks

NIC Legend:

  NIC0: mlx5_0
  NIC1: mlx5_1
  NIC2: mlx5_2
  NIC3: mlx5_3
  NIC4: mlx5_4
  NIC5: mlx5_5
  NIC6: mlx5_6


Hypervisor vendor:: KVM
ulimit soft: 1024

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions