[fix] restore expert_bias to fp32 before bridge weight export by yueming-yuan · Pull Request #811 · radixark/miles

yueming-yuan · 2026-03-26T23:23:01Z

Summary

Float16Module casts all buffers to bf16, including expert_bias which must stay fp32
Megatron has _maintain_float32_expert_bias() to restore it after forward(), but bridge export happens outside the forward path
This causes bf16-truncated expert_bias values to be sent to SGLang during weight update, failing --check-weight-update-equal
Fix: call _maintain_float32_expert_bias() in patch_megatron_model before bridge export reads the buffer

gemini-code-assist

Code Review

This pull request introduces logic to ensure that expert_bias remains in float32 format by calling _maintain_float32_expert_bias on relevant modules before bridge export. The reviewer suggested moving this logic, as well as the preceding configuration patching, inside the try block to ensure the finally block correctly handles cleanup if an exception occurs during the setup phase.

gemini-code-assist · 2026-03-26T23:24:59Z

+    # Float16Module casts buffers to bf16, but expert_bias must stay fp32.
+    # Restore before bridge export reads the values.
+    for m in model:
+        for module in m.modules():
+            if hasattr(module, "_maintain_float32_expert_bias"):
+                module._maintain_float32_expert_bias()
+
    try:
        yield


This new logic should be inside the try block. If an exception occurs here, the finally block won't be executed, leaving model_config in a patched state. This could lead to unexpected behavior.

For full robustness, the preceding if block that patches share_embeddings_and_output_weights should also be moved inside the try block.

Suggested change

# Float16Module casts buffers to bf16, but expert_bias must stay fp32.

# Restore before bridge export reads the values.

for m in model:

for module in m.modules():

if hasattr(module, "_maintain_float32_expert_bias"):

module._maintain_float32_expert_bias()

try:

yield

try:

# Float16Module casts buffers to bf16, but expert_bias must stay fp32.

# Restore before bridge export reads the values.

for m in model:

for module in m.modules():

if hasattr(module, "_maintain_float32_expert_bias"):

module._maintain_float32_expert_bias()

yield

Float16Module casts all buffers to bf16, including expert_bias which should stay fp32. Megatron's _maintain_float32_expert_bias() restores it after forward(), but bridge export happens outside forward path, reading bf16-truncated values. Call the restore before export.

fzyzcjy · 2026-03-27T14:11:18Z

+    # Restore before bridge export reads the values.
+    for m in model:
+        for module in m.modules():
+            if hasattr(module, "_maintain_float32_expert_bias"):


a bit hesitate about such hasattr operations, which may be error prone.

qq: do we have a CI to guarantee this is not regressed?

If the weight checker passs then it should be ok; I found this issue b/c the bias score cannot pass the weight checker in megatron bridge mode
maybe i can also add the 4 layer kimi k2.5 in CI and it will check

that looks great, adding ci then looks safe

fzyzcjy

LGTM if we add that to ci

…region clusters (#10) * Revert "[BUGFIX] [P2PRDMA] Add rollout post-processing after P2PRDMA weight updates" (radixark#882) * [Fix] fix ci (radixark#894) * Avoid threading for ray getting object (radixark#886) * Add explicit errors for unsupported Megatron profiles (radixark#887) * Add nvfp4 quantizer files (radixark#907) * Bump flash-linear-attention version to 0.4.2 (radixark#892) * [BUGFIX] Invoke "post_process_quantization" by default after weight updating (radixark#890) Co-authored-by: Yueming Yuan <yym022502@gmail.com> * Add heartbeat and id to session server (radixark#866) * fix: adding thin glm5 image to docker build + latest tag sync (radixark#871) * Add consistent hashing routing policy for rollout (radixark#891) Co-authored-by: Yueming Yuan <yueming@Mac.attlocal.net> * [example] add retool v2 example with multi-turn framework interfaces (radixark#654) Co-authored-by: GuanxingLu <gxlu02@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Expose rollout-batch-size, n-samples-per-prompt, global-batch-size as CLI args in swe-agent-v2 (radixark#954) Co-authored-by: Shi Dong <shi.dong@radixark.ai> * chore: remove obsolete swe-agent server.py and run-qwen3.sh (radixark#952) Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Add weight staleness control for fully async rollout (radixark#958) * Fix/pause generation mode (radixark#924) Co-authored-by: Yueming Yuan <yym022502@gmail.com> * [v0.5.10][1] Bump sglang to v0.5.10 (radixark#898) * [v0.5.10][2] Fix apply_chat_template behavior for transformers >=5.0 (radixark#926) Co-authored-by: guapisolo <guapisolo@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * [v0.5.10][3] Fix processor return_tensors duplicate kwarg for transformers >=5.0 (radixark#927) Co-authored-by: guapisolo <guapisolo@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * [v0.5.10][4] Fix _no_split_modules set not subscriptable in transformers >=5.0 (radixark#931) * [v0.5.10][5] Disable piecewise cuda graph to avoid NVLS oom (radixark#935) * [v0.5.10][6][FSDP] fix outdated weight update logic in FSDP (radixark#948) Co-authored-by: guapisolo <guapisolo@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: maocheng23 <35615230+maocheng23@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> * [v0.5.10][7][FSDP] move FSDP to experimental and disable by default (radixark#961) * Add skiplist and more robust calculation on val (radixark#965) * [fix] tiny fix debug rollout only in weight version check (radixark#967) * feat: real cp support with relayout fix for qwen3.5 train/rollout mismatch (radixark#885) * [AMD] Upgrade to sglv0.5.10 (radixark#973) * switch model to actor (radixark#756) * [fix] support general logic to bypass fp32 downcast and fix qwen35 A_log dtype (radixark#975) Co-authored-by: yueming-yuan <yym022502@gmail.com> * fix: populate prefix_cache_info in OpenAI/session rollout path (radixark#960) * Remove prepare_harbor_tasks.py; use harbor-private adapters (radixark#982) * [fix] Skip flush_cache in in_place mode and add fully async example (radixark#974) Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * GLM47 full cmd for async and sync reasoning (radixark#986) Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: handle non-tool appended messages in TITO incremental tokenization (radixark#949) Co-authored-by: Yanbin Jiang <jybsuper@gmail.com> * [docker] Add sgl-model-gateway install and download .tar.gz assets (radixark#895) * [ci] fix hf rate limit error by caching tokenizer loading (radixark#1014) Co-authored-by: maocheng23 <35615230+maocheng23@users.noreply.github.com> * Use load_generate_function in legacy sglang_rollout path (radixark#1016) * Update CODEOWNERS to add new reviewers (radixark#1021) * Support moe lora for gpt-oss (radixark#798) Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com> * [fix] restore expert_bias to fp32 before bridge weight export (radixark#811) * [chore] drop legacy transformers upgrade pin for glm47-flash and qwen35 (radixark#1018) * [fix] Enforce param dtype before wrap ddp (radixark#992) Co-authored-by: Zhichen Zeng <zczeng@uw.edu> * [upgrade] update Megatron-Bridge source and LoRA CI to megatron e2e tests and (radixark#1023) * [CI] Drop --use-miles-router from R3 tests and add r3 comparasion test between sgl & miles router (radixark#1015) * wandb: raise init_timeout, add retry wrapper, fix shared-mode init for cross-region clusters In online + shared mode, both `init_wandb_primary` and `init_wandb_secondary` make HTTPS round-trips to wandb cloud (login + run create/attach). On high-latency cross-region clusters (e.g. Abu Dhabi MBZUAI ↔ wandb-cloud US-West) with concurrent actor bursts, a single round-trip can exceed the wandb SDK's 90s default `init_timeout` — tearing down the whole run with a silent handshake abort. Observed on RL360 job 1564420, which forced `WANDB_MODE=offline` as a global default ever since (see https://github.com/LLM360/RL360/issues/87). The issue's original diagnosis assumed a local primary↔secondary socket handshake race. That's not how shared mode works — per wandb's own feature PR (wandb/wandb#6882), each writer spawns an independent wandb-core that talks to the cloud directly; aggregation is server-side by run_id. No local socket exists. The failure mode is pure network/latency, not a local readiness race. Changes ------- - Bump `init_timeout` to 300s for primary and secondary Settings. Configurable via `WANDB_INIT_TIMEOUT_SECS` env var for tuning. - Wrap both init paths in a bounded exponential-backoff retry (`_wandb_init_with_retry`) that re-attempts on wandb.errors.CommError and wandb.errors.UsageError. 3 attempts with 5→10→20s backoff by default, tunable via `WANDB_INIT_RETRY_ATTEMPTS` / `WANDB_INIT_RETRY_BACKOFF_SECS`. - Add `x_label` tagging per wandb distributed-training docs: primary gets `rank_<rank>_primary`, secondaries get `rank_<rank>_secondary`. Enables per-rank console-log filtering in the wandb UI. - Drop `reinit=True` from secondary init_kwargs. Shared mode natively supports concurrent writers on a single run; `reinit=True` triggered stale-state warnings on secondary actors without functional benefit. Followups this change enables ----------------------------- - `WANDB_MODE=offline` can be removed from scale.yaml's extra_env default once a pilot run confirms online mode boots cleanly. - The tmux-based `~/bin/wandb-sync-rl360.sh` workaround on David's M2 account becomes obsolete (no more offline-only default). - Near-realtime wandb dashboards replace the ~2-minute-lag offline sync; per-rank system metrics via x_label filtering. --------- Co-authored-by: JD <jaedon.guo@gmail.com> Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com> Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com> Co-authored-by: Ziang Li <ziangli@umich.edu> Co-authored-by: Zhichen Zeng <zczeng@uw.edu> Co-authored-by: JensenFire <xinji1@microsoft.com> Co-authored-by: Yueming Yuan <yym022502@gmail.com> Co-authored-by: maocheng23 <35615230+maocheng23@users.noreply.github.com> Co-authored-by: Douglas Yang <douglasyang88@gmail.com> Co-authored-by: Yueming Yuan <yueming@Mac.attlocal.net> Co-authored-by: Huapeng Zhou <73010314+PopSoda2002@users.noreply.github.com> Co-authored-by: GuanxingLu <gxlu02@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Shi-Dong <Shi-Dong@users.noreply.github.com> Co-authored-by: Shi Dong <shi.dong@radixark.ai> Co-authored-by: Jiajun Li <48857426+guapisolo@users.noreply.github.com> Co-authored-by: guapisolo <guapisolo@gmail.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: Yuzhen Zhou <82826991+zyzshishui@users.noreply.github.com> Co-authored-by: Yanbin Jiang <jybsuper@gmail.com> Co-authored-by: Ying Sheng <sqy1415@gmail.com> Co-authored-by: Yisheng Gong <yishenggong9437@gmail.com>

yueming-yuan requested review from fzyzcjy, guapisolo and maocheng23 as code owners March 26, 2026 23:23

gemini-code-assist Bot reviewed Mar 26, 2026

View reviewed changes

yueming-yuan changed the title ~~fix: restore expert_bias to fp32 before bridge weight export~~ [fix] restore expert_bias to fp32 before bridge weight export Mar 26, 2026

yueming-yuan force-pushed the fix/bridge-expert-bias-fp32 branch from e09ea02 to bb6e1c4 Compare March 26, 2026 23:35

fzyzcjy reviewed Mar 27, 2026

View reviewed changes

fzyzcjy approved these changes Mar 27, 2026

View reviewed changes

yueming-yuan merged commit 641f071 into main Apr 21, 2026
18 checks passed

yueming-yuan deleted the fix/bridge-expert-bias-fp32 branch April 21, 2026 05:13

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[fix] restore expert_bias to fp32 before bridge weight export#811

[fix] restore expert_bias to fp32 before bridge weight export#811
yueming-yuan merged 1 commit intomainfrom
fix/bridge-expert-bias-fp32

yueming-yuan commented Mar 26, 2026 •

edited

Loading

Uh oh!

gemini-code-assist Bot left a comment

Uh oh!

gemini-code-assist Bot Mar 26, 2026

Uh oh!

fzyzcjy Mar 27, 2026

Uh oh!

yueming-yuan Mar 27, 2026

Uh oh!

fzyzcjy Mar 27, 2026

Uh oh!

fzyzcjy left a comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Conversation

yueming-yuan commented Mar 26, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Summary

Uh oh!

gemini-code-assist Bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

gemini-code-assist Bot Mar 26, 2026

Choose a reason for hiding this comment

Uh oh!

fzyzcjy Mar 27, 2026

Choose a reason for hiding this comment

Uh oh!

yueming-yuan Mar 27, 2026

Choose a reason for hiding this comment

Uh oh!

fzyzcjy Mar 27, 2026

Choose a reason for hiding this comment

Uh oh!

fzyzcjy left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

yueming-yuan commented Mar 26, 2026 •

edited

Loading