Summary
Every tool that shells out to modelopt-launcher hangs until its internal timeout when modelopt-mcp is launched the way MCP clients launch it (stdio subprocess). The same command run manually from a shell finishes in ~2 s. One-line fix per call site: stdin=subprocess.DEVNULL.
Environment
|
|
| OS |
Windows 10 (x86_64) |
| Python |
3.12 (uv-managed venv, uv sync) |
| modelopt-mcp |
0.1.0 (main) |
| mcp |
1.30.0 (server env) |
| Client |
Hermes Agent, MCP SDK 2.0.0 (stdio transport) |
| GPU |
none usable (Quadro P2000, cc 6.1) — irrelevant, --dryrun needs no GPU |
Steps to reproduce
-
Sparse-checkout and install the server exactly as documented in tools/mcp/README.md:
git clone --depth 1 --filter=blob:none --sparse https://github.com/NVIDIA/Model-Optimizer.git
cd Model-Optimizer && git sparse-checkout set tools/mcp tools/launcher
cd tools/mcp && uv sync --python 3.12
-
Start the server as a stdio subprocess (any MCP client does this; .mcp.json in plugins/modelopt/ is exactly this shape) and call:
{"name": "submit_job",
"arguments": {"yaml_path": "examples/smoke/nvidia_smi.yaml", "dry_run": true}}
-
Observe: returns after 60 s with
{"ok": false, "dry_run": true, "reason": "dry_run_timeout", "exit_code": null,
"stdout_tail": "", "stderr_tail": "",
"diagnostic": "launch.py --dryrun did not return within 60 seconds..."}
Note stdout_tail/stderr_tail are empty — not even the resolved-job-spec table the launcher normally prints.
-
Control — run the identical argv by hand from a shell:
modelopt-launcher --yaml <abs>/examples/smoke/nvidia_smi.yaml --dryrun --yes
→ prints the full resolved spec and exits rc=0 in ~2.1 s.
Root cause
tools/mcp/modelopt_mcp/bridge.py spawns the launcher with capture_output=True, which redirects stdout/stderr only. stdin is left inherited. For a stdio MCP server, the inherited stdin is the MCP JSON-RPC transport pipe, so the child contends with the server for the same stdin handle and blocks.
Not specific to dry-run — all three launcher call sites are affected:
| Line |
Call |
Path affected |
| ~1257 |
subprocess.Popen |
Docker live submit |
| ~1298 |
subprocess.run (timeout=300) |
Slurm live submit |
| ~1507 |
subprocess.run (timeout=60) |
dry_run |
The comments at ~1454 already note that without --yes nemo_run's entrypoint blocks on a confirmation prompt "and since we're capturing stdout (no TTY), the prompt would hang until the 60-second timeout fires" — this is the same class of problem via a different handle.
Fix
subprocess.run(argv, env=child_env, capture_output=True, text=True,
timeout=60, check=False,
stdin=subprocess.DEVNULL) # dry_run
# and stdin=subprocess.DEVNULL on the Popen and the 300 s run as well
Verification after the fix
|
before |
after |
submit_job(dry_run=True) over stdio MCP |
dry_run_timeout @ 60 s, empty stdout |
{"ok": true, "dry_run": true, "validated": true, "exit_code": 0} in 2.2 s |
stdout_tail now contains the resolved SandboxTask0 spec, i.e. the YAML loader + factory resolution actually ran.
Notes
- A standalone repro with a fresh, never-written pipe (
subprocess.PIPE held by a plain parent) does not hang, so the trigger is specifically inheriting the server's live stdio pipe. Environment matrix (cwd × NEMORUN_HOME × stdin shape) was tested; only the MCP-server context reproduces it.
verify_setup is unaffected (its subprocesses fail fast with FileNotFoundError on hosts lacking docker/slurm); list_examples is pure Python.
- Suggested test: a
tests/-level regression that asserts every subprocess.* call in bridge.py passes an explicit stdin=.
Summary
Every tool that shells out to
modelopt-launcherhangs until its internal timeout whenmodelopt-mcpis launched the way MCP clients launch it (stdio subprocess). The same command run manually from a shell finishes in ~2 s. One-line fix per call site:stdin=subprocess.DEVNULL.Environment
uv sync)--dryrunneeds no GPUSteps to reproduce
Sparse-checkout and install the server exactly as documented in
tools/mcp/README.md:Start the server as a stdio subprocess (any MCP client does this;
.mcp.jsoninplugins/modelopt/is exactly this shape) and call:{"name": "submit_job", "arguments": {"yaml_path": "examples/smoke/nvidia_smi.yaml", "dry_run": true}}Observe: returns after 60 s with
{"ok": false, "dry_run": true, "reason": "dry_run_timeout", "exit_code": null, "stdout_tail": "", "stderr_tail": "", "diagnostic": "launch.py --dryrun did not return within 60 seconds..."}Note
stdout_tail/stderr_tailare empty — not even the resolved-job-spec table the launcher normally prints.Control — run the identical argv by hand from a shell:
→ prints the full resolved spec and exits rc=0 in ~2.1 s.
Root cause
tools/mcp/modelopt_mcp/bridge.pyspawns the launcher withcapture_output=True, which redirects stdout/stderr only.stdinis left inherited. For a stdio MCP server, the inherited stdin is the MCP JSON-RPC transport pipe, so the child contends with the server for the same stdin handle and blocks.Not specific to dry-run — all three launcher call sites are affected:
subprocess.Popensubprocess.run(timeout=300)subprocess.run(timeout=60)dry_runThe comments at ~1454 already note that without
--yesnemo_run'sentrypointblocks on a confirmation prompt "and since we're capturing stdout (no TTY), the prompt would hang until the 60-second timeout fires" — this is the same class of problem via a different handle.Fix
Verification after the fix
submit_job(dry_run=True)over stdio MCPdry_run_timeout@ 60 s, empty stdout{"ok": true, "dry_run": true, "validated": true, "exit_code": 0}in 2.2 sstdout_tailnow contains the resolvedSandboxTask0spec, i.e. the YAML loader + factory resolution actually ran.Notes
subprocess.PIPEheld by a plain parent) does not hang, so the trigger is specifically inheriting the server's live stdio pipe. Environment matrix (cwd ×NEMORUN_HOME× stdin shape) was tested; only the MCP-server context reproduces it.verify_setupis unaffected (its subprocesses fail fast withFileNotFoundErroron hosts lacking docker/slurm);list_examplesis pure Python.tests/-level regression that asserts everysubprocess.*call inbridge.pypasses an explicitstdin=.