Troubleshooting

Troubleshooting

Common problems when starting and running profiles, with the symptom, the underlying cause, and the fix. If a container fails, the fastest first step is always to open the profile in the TUI and stream its logs (l) — llmux already surfaces the last log lines on a failed start, but the full log tells the rest of the story.

Container stuck at health: starting

Symptom. The dashboard shows the profile as starting and it never moves to healthy. docker ps shows (health: starting).

Cause. This is almost always a model still downloading or loading. The health checks only pass once the server is actually serving:

A first start that downloads weights — a multi-GB HF repo, or a GGUF pulled by llama-server — can easily exceed those windows on a slow link.

Fix. Stream the logs and watch the download/load progress. Once the model finishes loading, the status flips to healthy on the next refresh. Subsequent starts reuse the cached files and come up fast.

GPU out of memory

Symptom. The container exits or goes unhealthy seconds after start. llmux reports container exited during startup with the last log lines, which mention CUDA OOM during engine/model init.

Cause. The model plus its KV cache does not fit in the allocated VRAM.

Fix.

  • Lower gpu-memory-utilization in config/vllm/<name>.yaml (it is the fraction of each GPU's VRAM vLLM may use).
  • Lower max-model-len — a shorter context length means a smaller KV cache.
  • Raise tensor_parallel_size on the profile (and add GPUs to gpu_id, e.g. "0,1") to shard the model across more GPUs.
  • Lower n-gpu-layers in config/llamacpp/<name>.yaml to offload more layers to CPU.
  • Quantize the KV cache with cache-type-k / cache-type-v (e.g. q8_0).
  • For MoE models, add an override-tensors rule (e.g. ".ffn_.*_exps.=CPU") to push expert tensors to system RAM.

Size it before you start it

Use the memory estimator (m in the TUI, or llmux system mem-estimate <model>) to check whether a model fits before launching it.

Port conflict

Symptom. The start is refused immediately with Error: Port <n> is already in use by ....

Cause. llmux checks the profile's port before starting. The port may be held by:

Fix. Change the profile's port in profiles.yaml (or via llmux profile edit <name> --port <n>), or stop whatever is holding the port. The cross-backend conflict gate also warns pre-start when another profile — on either backend — overlaps.

.env.common not found

Symptom. The start fails with Error: .env.common not found (vLLM) or ✗ .env.common 없음 (llama.cpp).

Cause. .env.common holds shared host paths and tokens and is gitignored — it does not exist until you create it.

Fix.

cp .env.common.example .env.common
$EDITOR .env.common      # set HF_CACHE_PATH, MODEL_DIR, HF_TOKEN, ...

Then run llmux env-check to confirm. Note that HF_CACHE_PATH must be an absolute path — a relative value is rejected. If the profile uses LoRA, LORA_BASE_PATH must also be set and absolute.

Gated model download fails (missing HF_TOKEN)

Symptom. The container starts but the engine cannot fetch the weights; logs show a 401/403 from Hugging Face, or "gated repo" / "access denied".

Cause. HF_TOKEN in .env.common is empty or lacks access to the repo. Public models download fine without a token, but gated repos (some Llama and Gemma weights, for example) do not.

Fix. Set HF_TOKEN in .env.common to a token with access to that repo. It is passed into both backends' containers (HUGGING_FACE_HUB_TOKEN for vLLM, HF_TOKEN for llama.cpp), so the in-container download can authenticate. Make sure you have also accepted the model's license on the Hugging Face website.

Dev image not found — "build it first"

Symptom. Starting a profile that points at a vllm-dev:<tag> or llamacpp-dev:<tag> image fails with Image ... not found and, for llama.cpp, a hint to build it.

Cause. Dev images are built locally and are not in any registry. The tag either was never built, or — for the auto-build paths — its repo/branch labels do not match the requested source.

Fix.

Swapping the source repo

If .vllm-src/ or .llamacpp-src/ already exists with a different origin URL, llmux refuses to switch remotes silently. Move or delete that checkout directory yourself, then rebuild.

Local Latest shows "no images" (vLLM)

Symptom. The vLLM start dialog's Local Latest option is labeled "no images" and the start is refused.

Cause. llmux only counts semver-style tags (vX.Y.Z) as "local latest". latest and nightly are deliberately skipped because they do not describe their own contents. If you only have those aliases locally, there is no versioned tag to pick.

Fix. Pull a specific version, then retry:

docker pull vllm/vllm-openai:v0.19.1

Or choose Official Release in the dialog, which resolves Docker Hub's current stable version and pulls it by explicit tag.

No GPU / nvidia-smi unavailable

Symptom. A dev build logs that it could not detect the GPU and falls back to a multi-arch build; or starting any profile fails because the GPU runtime is not visible.

Cause. The host has no working nvidia-smi, or Docker's NVIDIA Container Toolkit is not set up.

Fix.

Corrupted or incomplete model file

Symptom. The container starts but the engine crashes while loading the weights — checksum mismatch, truncated file, "failed to load model", or a GGUF/safetensors parse error.

Cause. A download that was interrupted (network drop, killed mid-pull) can leave a partial file in the Hugging Face cache.

Fix. Delete the bad file/blob from HF_CACHE_PATH (the host directory mounted into the container) and start again — both backends re-download anything missing from the cache. If it keeps failing on the same file, the upstream repo file itself may be the problem; try a different quantization or a mirror repo.

Copying logs out of the TUI

Symptom. You want to paste a container log somewhere but the TUI captures the keyboard.

Fix. Shift+drag to select text, then Ctrl+C to copy — this is standard Textual behavior. See the Textual FAQ.

See also