Guide Profiles Model Configs Container Lifecycle Using the TUI Dev Builds

Model Configs

A config is a YAML file holding the engine-level flags for one model. While a profile decides where a container runs (port, GPU, container name), the config decides how the model is served — context length, memory budget, sampling defaults, and so on.

Configs live under config/<backend>/<name>.yaml:

  • config/vllm/<name>.yaml — flags for vllm serve
  • config/llamacpp/<name>.yaml — flags for llama-server

Each profile links to one config via its config_name field (defaulting to the profile name). The key principle is 1:1 mapping: every key in a config maps directly to an engine command-line flag.

example.yaml is a template

Each backend ships a config/<backend>/example.yaml template. Copy it to a new name — cp config/vllm/example.yaml config/vllm/my-model.yaml — or use llmux config new. The example stem is excluded from config listings.

vLLM configs

Keys map 1:1 to vllm serve engine arguments. Use the long-flag name with -- stripped: --gpu-memory-utilization 0.9 becomes gpu-memory-utilization: 0.9. The full flag set is documented in the vLLM engine args reference.

# config/vllm/example.yaml
model: Qwen/Qwen3-8B               # HuggingFace repo id or local path

# ---- memory / parallelism ----
gpu-memory-utilization: 0.90       # 0.0-1.0, fraction of each GPU's VRAM vLLM may use
max-model-len: 32768               # context length (lower than model default = less memory)
# tensor-parallel-size: 2          # multi-GPU (keep in sync with the profile's GPU_ID)

# ---- sampling defaults (overridden per-request) ----
# dtype: bfloat16
# trust-remote-code: true
KeyTypeDescription
modelstringHugging Face repo id or local path. Required unless the profile sets model_id.
gpu-memory-utilizationfloatFraction (0.0-1.0) of each GPU's VRAM vLLM may claim.
max-model-lenintMaximum context length. Set below the model default to save memory.
tensor-parallel-sizeintTensor-parallel degree for multi-GPU. Keep consistent with the profile's gpu_id / tensor_parallel_size.
dtypestringWeight dtype, e.g. bfloat16.
trust-remote-codeboolAllow custom model code from the repo.

Any other vllm serve flag is valid here — these are just the common ones.

Where the config is mounted

The whole config/vllm/ directory is mounted into the container at /config/, and vLLM is launched with --config /config/<config_name>.yaml. The profile's tensor_parallel_size and LoRA options are passed as separate compose-level flags, not from this file.

llama.cpp configs

Keys map 1:1 to llama-server long flags: --ctx-size 32768 becomes ctx-size: 32768, --jinja becomes jinja: true, --temp 0.6 becomes temp: 0.6. At launch, render-override.py reads the config and renders a command: block for the container. See Container Lifecycle for that flow.

# config/llamacpp/example.yaml
model-file: your-model.gguf        # GGUF filename (relative to models/)

ctx-size: 8192                     # context length
n-gpu-layers: 99                   # layers to offload to GPU. 99 = all

cache-type-k: bf16                 # KV cache K precision: f16 / bf16 / q8_0 / q4_0 ...
cache-type-v: bf16                 # KV cache V precision

temp: 0.7                          # sampling defaults
top-p: 0.9
top-k: 40

alias: your-model                  # name shown in /v1/models

jinja: true                        # Jinja chat template (recommended for Qwen/Llama3 etc.)

# MoE expert CPU offload (for A3B-style models):
# regex-match tensor names → offload matched tensors to CPU
# override-tensors:
#   - ".ffn_.*_exps.=CPU"

# api-key: sk-your-key             # optional auth

extra-args: []                     # verbatim args render-override.py appends as-is
KeyTypeDescription
model-filestringGGUF filename. Acts as a display/fallback value — the actual download source is the profile's hf_repo / hf_file.
aliasstringModel name exposed at /v1/models.
ctx-sizeintContext length in tokens.
n-gpu-layersintNumber of layers offloaded to GPU. 99 = all.
cache-type-k / cache-type-vstringKV cache precision: f16, bf16, q8_0, q4_0, etc.
temp / top-p / top-k / min-pfloat / intDefault sampling parameters (overridable per request).
jinjaboolUse the model's Jinja chat template. Recommended for modern models.
flash-attnboolEnable Flash Attention.
parallelintNumber of concurrent decoding slots.
cont-batchingboolContinuous batching.
metricsboolExpose the /metrics Prometheus endpoint.
api-keystringRequire this key on incoming requests.
override-tensorslistMoE CPU offload. Each entry is a <regex>=<device> rule, e.g. ".ffn_.*_exps.=CPU", applied as -ot. Keeps experts on CPU while attention/embeddings stay on GPU.
extra-argslistEscape hatch — items are appended to the llama-server command verbatim, for newer flags not expressible in the schema above.

host / port / no-webui are forced

render-override.py always sets --host 0.0.0.0 --port 8080 --no-webui and drops any host, port, webui, or no-webui keys from your config. Map the host port via the profile's port field instead.

MoE expert CPU offload

For Mixture-of-Experts models (e.g. Qwen3-A3B), override-tensors lets you keep the large expert tensors on CPU while attention and embeddings stay on GPU:

n-gpu-layers: 99
override-tensors:
  - ".ffn_.*_exps.=CPU"   # FFN expert tensors → CPU; everything else → GPU

Each rule renders to a -ot <rule> argument.

Import a vLLM recipe

vLLM publishes verified serving configs per model at vllm-project/recipes (rendered on recipes.vllm.ai). llmux can pull one and turn it into a config — but a recipe verified on an 80 GB card may not fit yours, so it never applies blindly.

In the vLLM Quick Setup form, type a model id and press 📋 Fetch vLLM recipe. A review window opens showing:

Confirm and it creates the profile + config (swapping in the FP8/AWQ checkpoint and quantization flags when you pick that variant).

Headless equivalent:

# See the recipe's variants + features
llmux config from-recipe Qwen/Qwen3-32B --list

# Auto-pick a GPU-fitting variant (or force one with --variant)
llmux config from-recipe Qwen/Qwen3-32B --variant fp8 --feature reasoning

# Preview without writing
llmux config from-recipe Qwen/Qwen3-32B --variant awq --json

Flag autocomplete, from your actual image

When you add a flag row in the TUI config form, the name field autocompletes against the real flag set of the engine you're running — not a hardcoded list. The first time it's needed, llmux runs your vLLM image once (vllm serve --help) or llama-server --help, extracts every valid flag, and caches the result per image tag. Pin a different vLLM version to a profile and its config form suggests that version's flags.

A one-line description of the focused flag is shown underneath the row, so you can check what a flag does without leaving the form.

Editing configs from the CLI

# List configs across both backends
llmux config list

# Print one config
llmux config show my-model

# Create a vLLM config
llmux config new my-model --backend vllm --model Qwen/Qwen3-8B --gpu-mem 0.9

# Create a llama.cpp config, setting flags inline
llmux config new gemma --backend llamacpp \
    --set model-file=gemma-3-4b-it-q4_k_m.gguf \
    --set ctx-size=8192 \
    --set n-gpu-layers=99 \
    --set jinja=true

# Patch an existing config
llmux config edit gemma --set temp=0.6 --set top-p=0.95
llmux config edit gemma --unset top-k

# Clone a tuned config to a new name
llmux config clone gemma gemma-longctx

# Rename in place — profiles pointing at it are repointed automatically
llmux config rename gemma gemma-q4

# Delete
llmux config delete gemma

--set values are YAML-typed

--set parses each value as YAML: --set ctx-size=8192 stores an int, --set jinja=true stores a bool, and an empty value (--set jinja=) stores true. For list-valued keys like override-tensors or extra-args, pass a YAML literal — --set 'override-tensors=[".ffn_.*_exps.=CPU"]' — or edit the file directly / use the TUI config form. A plain string works too: --set 'override-tensors=.ffn_.*_exps.=CPU'.

Renaming a config

A config's name is its filename (config/<backend>/<name>.yaml), and profiles reference it by that name. config rename moves the file and rewrites every referencing profile's config_name in one step, so nothing is left pointing at a file that no longer exists.

llmux config rename gemma gemma-q4
# Renamed config/llamacpp/gemma.yaml → config/llamacpp/gemma-q4.yaml
# Repointed profile: gemma

In the TUI, press R on the config list. A rename is refused while any referencing profile's container is running — that container would keep serving the old file until it restarts, leaving the running server disagreeing with profiles.yaml.

A profile that never set config_name resolves to a config matching its own name, and those implicit references are repointed too. The tracked example.yaml template cannot be renamed.

Toggling parameters on / off

Sometimes you want to keep a flag but stop passing it to the server — an A/B test, a setting that only applies to one model variant, a tweak you're not sure about. Instead of deleting the line (and losing its value), toggle it:

# Park a flag without deleting it
llmux config edit exaone-off --disable reasoning-parser

# Turn it back on later — the exact value comes back
llmux config edit exaone-off --enable reasoning-parser

A disabled parameter is stored as an inert comment marker at the end of the file:

model: LGAI-EXAONE/EXAONE-4.0-1.2B
gpu-memory-utilization: 0.6
# llmux:disabled reasoning-parser: deepseek_r1

Because it's a comment, neither vllm serve (which reads the YAML directly) nor llama.cpp's flag renderer ever sees it — but llmux does, so it can flip it back on. In the TUI config form each parameter row has an on/off switch; a switched-off row dims but keeps its value. --set KEY=value re-enables a disabled key automatically; --unset KEY removes it from both states.

Your comments survive edits

Editing a config through the TUI or CLI preserves any hand-written # comments — headers, inline notes, trailing explanations — for every parameter that survives the edit. Comment-free files are left byte-for-byte unchanged.

See also