Repository navigation
Sm90 mega moe on sgl dev - #36
Conversation
|
@qiushixiaoyu Can we upstream this change to original DeepGemm? So that we can use it more conveniently in the future |
@Fridge003 |
|
When enable --run-low-latency-baseline, Will there be a performance degradation? |
067fc03 to
78772d1
Compare
78772d1 to
fce68b3
Compare
I don’t think so. This is only for comparing performance against the low-latency baseline. While testing, I found that the performance with small batch sizes is not very stable. I’m still investigating it. |
|
@qiushixiaoyu Can you share your sglang start cmd? I try use this PR, but sglang output is error。 SGLANG_DSV4_FP4_EXPERTS=0 \
sglang serve \
--trust-remote-code \
--model-path /data1/DeepSeek-V4-Flash-FP8 \
--tp 8 \
--moe-a2a-backend megamoe \
--tool-call-parser deepseekv4 \
--reasoning-parser deepseek-v4 \
--host 0.0.0.0 \
--port 8055 |
@yz-tang I still have an SGLang change PR that hasn’t been merged yet. |
Co-authored-by: yinding <yinding@bytedance.com>
|
Hi,Have you verified the correctness with a max_token_per_rank=8K? I observed a significant drop in AIME results when running AIME25 with a concurrency of 40,max_token_per_rank=8K(sgl-version: 0.5.15,run deepseek v4 flash)Or am I using it incorrectly somewhere? |
Co-authored-by: yinding <yinding@bytedance.com>
Ports upstream PR sgl-project#36 ("Sm90 mega moe on sgl dev"), which added a Hopper FP8xFP8 fused MoE GEMM kernel on the old `dev` branch layout (tvm_ffi_api.cpp / sgl_deep_gemm), onto nv_dev's current layout (csrc/apis/*.hpp, csrc/python_api.cpp, deep_gemm/mega). nv_dev already had MegaMoE, but only for SM100 (Blackwell), built on a ring-buffer / cluster-pair-interleaved persistent-grid scheduler and a UE8M0/FP4-oriented workspace layout. The SM90 kernel in PR deepseek-ai#36 instead uses a simpler pool-based workspace and per-expert-wave scheduler with float (non-UE8M0) scale factors, FP8-only weights (no FP4), and no 2-CTA clusters or shared-expert support -- these are not compatible data layouts, so the SM90 path is added as a fully parallel path alongside (not replacing) the SM100 one. New files (ported from the PR's `.cuh`/`.hpp`, adapted to nv_dev's current naming/dispatch conventions and symbol-renamed to avoid colliding with SM100's `layout::Workspace` / `sched::MegaMoEScheduler` / `sched::BlockPhase`): - deep_gemm/include/deep_gemm/layout/sm90_mega_moe.cuh (`layout::MegaMoESM90Workspace`, pool-based; reuses the existing shared `layout::Data`/`layout::Buffer`/`layout::TokenSrcMetadata` and `layout::get_num_max_pool_tokens` from layout/mega_moe.cuh) - deep_gemm/include/deep_gemm/scheduler/sm90_mega_moe.cuh (`sched::MegaMoESM90Scheduler` / `sched::MegaMoESM90BlockPhase`) - deep_gemm/include/deep_gemm/impls/sm90_fp8_mega_moe.cuh (the ported kernel, ~1935 lines, WGMMA/TMA-based Hopper impl) - csrc/jit_kernels/heuristics/sm90_mega_moe.hpp (`MegaMoESM90Config` and block/pipeline/wave heuristics, mirroring csrc/jit_kernels/heuristics/mega_moe.hpp's structure) - csrc/jit_kernels/impls/sm90_fp8_mega_moe.hpp (JIT host runtime, mirroring csrc/jit_kernels/impls/sm100_fp8_fp4_mega_moe.hpp) - tests/test_mega_moe_hopper.py (correctness test with an SM90 capability guard; skips cleanly on non-Hopper GPUs) Modified files: - csrc/apis/mega.hpp: adds `get_symm_buffer_size_for_sm90_mega_moe` and `fp8_mega_moe_sm90`, registered via the existing `deep_gemm::mega::register_apis` (already wired into csrc/python_api.cpp, so no python_api.cpp changes were needed). - deep_gemm/mega/__init__.py + deep_gemm/__init__.py: expose `Sm90SymmBuffer`, `get_symm_buffer_for_sm90_mega_moe`, `transform_weights_for_mega_moe_sm90`, `fp8_mega_moe_sm90`, following the existing SM100 exposure pattern. - deep_gemm/include/deep_gemm/comm/barrier.cuh: generalizes `grid_sync`/`nvlink_barrier` to a templated workspace type (instead of hard-coding `layout::Workspace`) so the SM90 kernel can reuse them with `layout::MegaMoESM90Workspace`; existing SM100 call sites are unaffected (the type is still deduced from the argument). Also ports the PR's ARCH 900-1000 guarded trap-instead -of-printf change to the NVLink barrier timeout path. Not ported: shared-expert support and the `situ` activation (the PR does not implement either for SM90). No GPU/CUDA toolchain is available in this environment, so this is a best-effort, close-reading port verified via `python3 -m py_compile`, brace/paren balance checks on all new/modified C++/CUDA files, and confirming `import deep_gemm` fails identically (missing compiled `_C` extension) before and after these changes -- i.e. no regression introduced. It has not been compiled or run on real Hopper hardware. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
MegaMoE SM90 Perf Summary
Flash vs normal baseline
Flash vs low-latency baseline
Pro vs normal baseline
Pro vs low-latency baseline
Benchmark DeepSeekV4Flash
OP Accuracy
correctness:28 scenarios PASS,max diff 0.0006
E2E Accuracy