-
Notifications
You must be signed in to change notification settings - Fork 727
Fix pruned-HF export fallback + add Nemotron-3.5-Lightning launcher examples #2196
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
422e5b0
e6e1fd6
0bd8609
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,56 @@ | ||
| # Nemotron-3.5-Lightning-30B-A3B (MoE) pruning to 3B active via Megatron-Bridge (4 GPUs), then vLLM gen. | ||
| # | ||
| # Slurm: uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_prune.yaml --yes | ||
| # Local: uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_prune.yaml hf_local=/mnt/hf-local --yes | ||
|
|
||
| # NOTE: sized for fast CI; bump for production, e.g. --calib_num_samples 1024 --seq_length 8192 --top_k 10. | ||
| # May need to reduce batch size if running out of memory at large seq_length. | ||
| job_name: Nemotron-3.5-Lightning-30B-A3B_mbridge_prune | ||
| pipeline: | ||
| note: "Prune Nemotron-3.5-Lightning-30B-A3B -> 3B active (Megatron-Bridge) with MMLU gate, then vLLM gen" | ||
|
|
||
| global_vars: | ||
| # Per-run scratch (fresh cicd_<id> dir) so each run prunes fresh | ||
| output_dir: /scratchspace/Nemotron-3.5-Lightning-30B-A3B-Pruned-A3.0B | ||
|
|
||
| # 1) Prune and export as a HF checkpoint. | ||
| # --score_lower_bound fails the job if the pruned model's MMLU drops below the floor. | ||
| task_0: | ||
| environment: | ||
| - LAUNCH_SCRIPT: torchrun --nproc_per_node 4 | ||
| inline: >- | ||
| $LAUNCH_SCRIPT modules/Model-Optimizer/examples/megatron_bridge/prune_minitron.py | ||
| --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | ||
| --trust_remote_code | ||
| --pp_size 4 | ||
| --calib_batch_size 8 | ||
| --calib_num_samples 256 | ||
| --seq_length 512 | ||
| --prune_target_active_params 3e9 | ||
| --prune_target_params 24e9 | ||
| --prune_score_func mmlu_10pct_bs32 | ||
| --max_width_pruning 0.30 | ||
| --max_depth_pruning 0.15 | ||
| --hparams_to_skip num_attention_heads | ||
| --top_k 5 | ||
| --score_lower_bound 0.58 | ||
| --output_hf_path <<global_vars.output_dir>> | ||
| slurm_config: &sc | ||
| _factory_: "slurm_factory" | ||
| container: nvcr.io/nvidia/nemo:26.08 | ||
| modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt | ||
| docker_user: root | ||
| nodes: 1 | ||
| ntasks_per_node: 4 | ||
| gpus_per_node: 4 | ||
|
kevalmorabia97 marked this conversation as resolved.
|
||
|
|
||
| # 2) vLLM sanity generation on the pruned checkpoint. | ||
| task_1: | ||
| inline: >- | ||
| python modules/Model-Optimizer/examples/megatron_bridge/generate_vllm.py | ||
| --model <<global_vars.output_dir>> | ||
| --trust_remote_code | ||
| --tensor_parallel_size 4 | ||
| slurm_config: | ||
| <<: *sc | ||
| ntasks_per_node: 1 | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,62 @@ | ||
| # Nemotron-3.5-Lightning-30B-A3B NVFP4 (W4A16 4/6) quantization + unified-HF export via Megatron-Bridge (4 GPUs). | ||
| # | ||
| # Slurm: uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml --yes | ||
| # Local: uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml hf_local=/mnt/hf-local --yes | ||
|
coderabbitai[bot] marked this conversation as resolved.
|
||
|
|
||
| # NOTE: sized for fast run; bump for production, e.g. --calib_num_samples 512 --seq_length 8192. | ||
| job_name: Nemotron-3.5-Lightning-30B-A3B_mbridge_quantize | ||
| pipeline: | ||
| note: "NVFP4 W4A16 PTQ Nemotron-3.5-Lightning-30B-A3B (Megatron-Bridge recipe), then unified-HF export" | ||
|
|
||
| global_vars: | ||
| # Per-run scratch (fresh cicd_<id> dir) so each run quantizes fresh. | ||
| output_dir: /scratchspace | ||
|
|
||
| # 1) NVFP4 Quantize via the ptq recipe and export to a deployable unified-HF checkpoint. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. since there's an existing mbridge_qad.yaml example that includes PTQ & Export is it possible to reuse that but skip QAD?
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I want to add this as a nmm-sandbox test to catch regressions. QAD test requires more compute so not sure if it will be enabled in sandbox tests or not
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. could you still reuse the qad example but add a variable to enable skipping QAD?
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. you could even add MMLU as a 4th step there but make it optional
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. QAD script is configured for 32k seq len, 512 samples while the quantize one I am using 512 seq len and 256 samples to make it run faster. Would have to
It might require more changes in our launcher core logic to support this. I will leave it out of this PR |
||
| # tp_size=1: static-block NVFP4 (MSE) weight quant is unsupported with TP>1. | ||
| task_0: | ||
| environment: | ||
| - LAUNCH_SCRIPT: torchrun --nproc_per_node 4 | ||
| inline: >- | ||
| $LAUNCH_SCRIPT modules/Model-Optimizer/examples/megatron_bridge/quantize.py | ||
| --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | ||
| --trust_remote_code | ||
| --tp_size 1 | ||
| --recipe huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6 | ||
| --calib_batch_size 8 | ||
| --calib_num_samples 256 | ||
| --seq_length 512 | ||
| --skip_generate | ||
| --export_megatron_path <<global_vars.output_dir>>/Nemotron-3.5-Lightning-30B-A3B-NVFP4-megatron | ||
| && | ||
| $LAUNCH_SCRIPT modules/Model-Optimizer/examples/megatron_bridge/export_quantized_megatron_to_hf.py | ||
| --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | ||
| --megatron_path <<global_vars.output_dir>>/Nemotron-3.5-Lightning-30B-A3B-NVFP4-megatron | ||
| --trust_remote_code | ||
| --pp_size 4 | ||
| --export_unified_hf_path <<global_vars.output_dir>>/Nemotron-3.5-Lightning-30B-A3B-NVFP4-hf | ||
| slurm_config: &sc | ||
| _factory_: "slurm_factory" | ||
| container: nvcr.io/nvidia/nemo:26.08 | ||
| modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt | ||
| docker_user: root | ||
| nodes: 1 | ||
| ntasks_per_node: 4 | ||
| gpus_per_node: 4 | ||
|
|
||
| # 2) MMLU (10% sample) on the exported NVFP4 checkpoint via vLLM, gated on a lower bound. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. is MMLU still a good metric to optimize for? maybe MMLU Pro or an agentic benchmark would be better suited
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Ideally yes we want to use new metrics but needs more work to set it up in sandbox tests. But currently MMLU is simplest to use using our lm eval scripts. Same approach as MLM sandbox tests. |
||
| task_1: | ||
| reqs_file: modules/Model-Optimizer/examples/llm_eval/requirements.txt | ||
| inline: >- | ||
| python modules/Model-Optimizer/examples/llm_eval/lm_eval_hf.py | ||
| --model vllm | ||
| --model_args pretrained=<<global_vars.output_dir>>/Nemotron-3.5-Lightning-30B-A3B-NVFP4-hf,tensor_parallel_size=4 | ||
| --trust_remote_code | ||
| --tasks mmlu | ||
| --limit 0.1 | ||
| --batch_size auto | ||
| --output_path /scratchspace/mmlu_results | ||
| --accuracy_lower_bound 0.75 | ||
| slurm_config: | ||
| <<: *sc | ||
| ntasks_per_node: 1 | ||
Uh oh!
There was an error while loading. Please reload this page.