-
Notifications
You must be signed in to change notification settings - Fork 247
[AMD] [WIP] [AGENTX] MiniMax-M3 Support on MI355X with MTP #2487
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
ajith-sirra-amd
wants to merge
5
commits into
main
Choose a base branch
from
amd/agentx-minimax-m3-mtp
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
+176
−0
Open
Changes from all commits
Commits
Show all changes
5 commits
Select commit
Hold shift + click to select a range
7906bed
[AMD] [WIP] [AGENTX] MiniMax-M3 Support on MI355X with MTP
AjithSirra 4edbe0e
[AMD] [WIP] [AGENTX] MiniMax-M3 Support on MI355X with MTP - Modifyin…
AjithSirra 635a5d6
[AMD] [WIP] [AGENTX] MiniMax-M3 Support on MI355X with MTP - Adding P…
AjithSirra f240481
[AMD] [WIP] [AGENTX] MiniMax-M3 Support on MI355X with MTP - Adding D…
AjithSirra cebeb27
feat(minimaxm3-fp4-mi355x-vllm-agentic-mtp): pin synthetic EAGLE3 acc…
functionstackx File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
154 changes: 154 additions & 0 deletions
154
benchmarks/single_node/agentic/minimaxm3_fp4_mi355x_mtp.sh
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,154 @@ | ||
| #!/usr/bin/env bash | ||
| set -euo pipefail | ||
| set -x | ||
|
|
||
| # Agentic trace replay benchmark for Minimax-M3 FP4 on MI355X using vLLM. | ||
| # | ||
| # Required env vars: | ||
| # MODEL, MODEL_PATH, TP, CONC, KV_OFFLOADING, KV_OFFLOAD_BACKEND, | ||
| # TOTAL_CPU_DRAM_GB, RESULT_DIR, DURATION, EP_SIZE, DP_ATTENTION | ||
|
|
||
| source "$(dirname "$0")/../../benchmark_lib.sh" | ||
|
|
||
| # Force the eval framework to lm-eval for this recipe. run_eval derives its | ||
| # default as swebench for agentic scenarios (scenario_default=swebench when | ||
| # IS_AGENTIC/SCENARIO_TYPE=agentic-coding), but EVAL_FRAMEWORK takes precedence | ||
| # over that default (benchmark_lib.sh: framework=${EVAL_FRAMEWORK:-...}), so | ||
| # setting it here makes the effective framework always lm-eval, never swebench. | ||
| export EVAL_FRAMEWORK="lm-eval" | ||
|
|
||
| check_env_vars MODEL TP CONC KV_OFFLOADING KV_OFFLOAD_BACKEND TOTAL_CPU_DRAM_GB RESULT_DIR DURATION EP_SIZE DP_ATTENTION | ||
|
|
||
| echo "MODEL=$MODEL TP=$TP CONC=$CONC KV_OFFLOADING=$KV_OFFLOADING TOTAL_CPU_DRAM_GB=$TOTAL_CPU_DRAM_GB RESULT_DIR=$RESULT_DIR DURATION=$DURATION EP_SIZE=$EP_SIZE DP_ATTENTION=$DP_ATTENTION" | ||
|
|
||
| PORT=8888 | ||
|
|
||
| if [[ -n "${SLURM_JOB_ID+x}" ]]; then | ||
| echo "JOB $SLURM_JOB_ID running on $SLURMD_NODENAME" | ||
| fi | ||
|
|
||
| # ROCR/HIP visibility for vLLM 0.14+ | ||
| if [[ -n "${ROCR_VISIBLE_DEVICES+x}" ]]; then | ||
| export HIP_VISIBLE_DEVICES="$ROCR_VISIBLE_DEVICES" | ||
| fi | ||
|
|
||
| if [[ -n "${MODEL_PATH:-}" ]]; then | ||
| if [[ ! -d "$MODEL_PATH" || -z "$(ls -A "$MODEL_PATH" 2>/dev/null)" ]]; then | ||
| hf download "$MODEL" --local-dir "$MODEL_PATH" | ||
| fi | ||
| else | ||
| hf download "$MODEL" | ||
| export MODEL_PATH="$MODEL" | ||
| fi | ||
|
|
||
| rocm-smi || true | ||
| amd-smi || true | ||
|
|
||
| resolve_trace_source | ||
| install_agentic_deps | ||
|
|
||
| DRAFT_MODEL="Inferact/MiniMax-M3-EAGLE3" | ||
| NUM_SPEC_TOKENS=5 | ||
| # Throughput pins synthetic EAGLE3 acceptance to the minimax-m3 golden AL | ||
| # (thinking_on, num_speculative_tokens=5, golden_al_distribution/minimaxm3_eagle3.yaml). | ||
| # MiniMax-M3 is an interleaved-thinking model and the replay runs it with | ||
| # thinking active, so thinking_on is the matching curve. The EVAL_ONLY accuracy | ||
| # run uses real target verification instead -- synthetic acceptance bypasses | ||
| # verification and corrupts the eval score. | ||
| SYNTHETIC_ACCEPT_LEN=3.35 | ||
| if [ "${EVAL_ONLY:-false}" = "true" ]; then | ||
| SPEC_CONFIG="{\"method\": \"eagle3\", \"model\": \"$DRAFT_MODEL\", \"num_speculative_tokens\": $NUM_SPEC_TOKENS, \"attention_backend\": \"TRITON_ATTN\"}" | ||
| else | ||
| SPEC_CONFIG="{\"method\": \"eagle3\", \"model\": \"$DRAFT_MODEL\", \"num_speculative_tokens\": $NUM_SPEC_TOKENS, \"attention_backend\": \"TRITON_ATTN\", \"rejection_sample_method\": \"synthetic\", \"synthetic_acceptance_length\": $SYNTHETIC_ACCEPT_LEN}" | ||
| fi | ||
|
|
||
| # ---- Server config ---------------------------------------------------------- | ||
| SERVER_LOG="$RESULT_DIR/server.log" | ||
| LMCACHE_LOG="$RESULT_DIR/lmcache_server.log" | ||
| mkdir -p "$RESULT_DIR" | ||
|
|
||
| OFFLOAD_ARGS=(--no-enable-prefix-caching) | ||
|
|
||
| case "$KV_OFFLOAD_BACKEND" in | ||
| vllm-simple) | ||
| unset VLLM_USE_SIMPLE_KV_OFFLOAD | ||
| # Use vLLM's regular native KV-offload path (OffloadingConnector), | ||
| # NOT the SimpleCPUOffloadConnector. The "native" backend resolves to | ||
| # OffloadingConnector by default; setting VLLM_USE_SIMPLE_KV_OFFLOAD=1 | ||
| # would switch it to SimpleCPUOffloadConnector. We intentionally leave | ||
| # that env var UNSET here so the regular OffloadingConnector path is | ||
| # used. The shortcut --kv_offloading_backend native + --kv_offloading_size | ||
| # form constructs the KVTransferConfig at engine startup | ||
| # (vllm/config/vllm.py:662). | ||
|
|
||
| # Remove --disable-hybrid-kv-cache-manager and enable hybrid kv cache manager (default) | ||
| # This gives extra cache hit than disabling hybrid kv cache manager | ||
| OFFLOAD_ARGS=( | ||
| --kv_offloading_backend native | ||
| --kv_offloading_size "$TOTAL_CPU_DRAM_GB" | ||
| ) | ||
| ;; | ||
| esac | ||
|
|
||
| # ---- LLM server config ---------------------------------------------------------- | ||
| PARALLEL_ARGS=(--tensor-parallel-size "$TP") | ||
| if [ "${DP_ATTENTION}" = "true" ]; then | ||
| PARALLEL_ARGS=( | ||
| --tensor-parallel-size 1 | ||
| --data-parallel-size "$TP" | ||
| --enable-expert-parallel | ||
| ) | ||
| elif [ "$EP_SIZE" -gt 1 ]; then | ||
| PARALLEL_ARGS+=(--enable-expert-parallel) | ||
| fi | ||
|
|
||
| echo "Starting vllm server..." | ||
| export PYTHONNOUSERSITE=1 | ||
|
|
||
| export VLLM_ENGINE_READY_TIMEOUT_S=3600 | ||
| export VLLM_USE_BREAKABLE_CUDAGRAPH=0 | ||
| export VLLM_ROCM_USE_AITER=1 | ||
| export VLLM_ROCM_USE_AITER_MOE=1 | ||
| export VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1 | ||
| # INT4 quantized all-reduce for the (~1.5 MB) decode all-reduces, which are the | ||
| # single biggest decode kernel at high concurrency. The MIN_SIZE_KB override is | ||
| # required: vLLM's default INT4 quick-reduce size gate for (bf16, TP4) is 16 MB, | ||
| # so it never fires for decode-sized tensors without it. | ||
| export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4 | ||
| export VLLM_ROCM_QUICK_REDUCE_CAST_BF16_TO_FP16=0 | ||
| export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION_MIN_SIZE_KB=256 | ||
|
|
||
| VLLM_CMD=( | ||
| vllm serve "$MODEL_PATH" | ||
| --served-model-name "$MODEL" | ||
| --host 0.0.0.0 | ||
| --port "$PORT" | ||
| "${PARALLEL_ARGS[@]}" | ||
| --speculative-config "$SPEC_CONFIG" | ||
| --trust-remote-code | ||
| --block-size 128 | ||
| --gpu-memory-utilization 0.80 | ||
| --language-model-only | ||
| --attention-backend TRITON_ATTN | ||
| --moe-backend aiter | ||
| --kv-cache-dtype fp8 | ||
| --tool-call-parser minimax_m3 | ||
| --enable-auto-tool-choice | ||
| --max-num-seqs "$CONC" | ||
|
ajith-sirra-amd marked this conversation as resolved.
|
||
| "${OFFLOAD_ARGS[@]}" | ||
| ) | ||
| printf '%q ' "${VLLM_CMD[@]}" | tee "$RESULT_DIR/vllm_command.txt" | ||
| printf '\n' | tee -a "$RESULT_DIR/vllm_command.txt" | ||
| "${VLLM_CMD[@]}" > "$SERVER_LOG" 2>&1 & | ||
| SERVER_PID=$! | ||
| echo "Server PID: $SERVER_PID" | ||
|
|
||
| wait_for_server_ready --port "$PORT" --server-log "$SERVER_LOG" --server-pid "$SERVER_PID" | ||
|
|
||
| # ---- Run benchmark ---------------------------------------------------------- | ||
| if [ "${EVAL_ONLY}" = "true" ]; then | ||
| run_eval --port "$PORT" | ||
| else | ||
| build_replay_cmd "$RESULT_DIR" | ||
| run_agentic_replay_and_write_outputs "$RESULT_DIR" | ||
| fi | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.