Two flags and a fortnight: getting Gemma 4 31B to fly on Intel Arc Pro B60s
Technical · Intel Arc Pro B60
Two Flags and a Fortnight: Getting Gemma 4 to Fly on Intel Arc Pro B60s
Eight Intel Arc Pro B60s, 191 GiB of VRAM, and a software stack that hides its worst failures behind fluent, plausible, wrong answers. Here is the configuration that works, the bugs that were costing us a factor of four, a vLLM defect that broke every small Gemma 4 model, and a like-for-like benchmark against an RTX PRO 6000 Blackwell at a fifth of the price.
Updated 1 October 2026. Since first publication we have retuned gpt-oss-120B, found and fixed a vLLM bug in the Gemma 4 E-series models, and re-run the NVIDIA comparison. What changed:
- Our earlier advice to cap
--max-num-seqsat 32 was wrong. The ceiling we measured was the cap itself. At 64, gpt-oss delivers 28–38% more. - gpt-oss now runs an fp8 KV cache, doubling its pool for about 1% of chat throughput. The matched NVIDIA comparison moves from 58% to 61% of the RTX PRO 6000.
- Expert parallelism was measured and rejected: 31% slower.
- New: Gemma 4 E4B was producing confident garbage on stock vLLM. We traced it to one misplaced line, and fixed it does 2,370 tokens per second on a single die.
- We now say plainly which of our settings depend on local patches to vLLM, and which work on Intel’s stock image.
The box
Four dual-die Arc Pro B60 cards, eight GPUs, 191 GiB of usable VRAM in a single-socket EPYC chassis. On paper it is a lot of memory bandwidth for the money. In practice, getting it to serve reliably took three weeks and turned up genuine defects in both the Intel vLLM stack and upstream vLLM — two of which silently corrupt output without failing anything.
We run it as two independent tensor-parallel-4 lanes: Gemma 4 31B for anything involving images on dies 0–3, and gpt-oss-120B for text and agent work on dies 4–7. They run simultaneously with no measurable interference. Small Gemma 4 models run one per die when we need them.
Single-stream decode, by configuration
The setting that matters most
If you take one thing from this article, take this one, because without it everything else you measure is fiction.
VLLM_XPU_INPLACE_ALLREDUCE=0vLLM’s XPU communicator performs an in-place all-reduce — it returns the input tensor as the output — and guards that with a single check:
# vllm/distributed/device_communicators/xpu_communicator.py
output = (
input_
if not _SKIP_ALL_REDUCE
and _INPLACE_ALL_REDUCE
and not torch.compiler.is_compiling()
else input_.clone()
)Elsewhere in the same codebase, parallel_state.py makes the equivalent decision using two conditions:
not torch.compiler.is_compiling()
and not torch.xpu.is_current_stream_capturing()During graph capture at runtime, is_compiling() returns False. So capture takes the aliasing path and bakes an aliasing all-reduce into the captured graph. At tensor-parallel 4 that corrupts every layer’s output. Torch even warns about it — “the output of this custom operator must not also be an input … please instead return a clone” — in a log line that scrolls past during startup.
The failure mode is the dangerous part. The model does not crash. It does not produce obvious garbage. It produces fluent, confident, plausible text that is wrong — and on image input, detailed descriptions of pictures it is not seeing. We benchmarked a configuration for an hour before discovering it had been answering incorrectly the whole time.
Controlled A/B on identical configuration, one variable: in-place, the correctness gate fails 2 of 3 cases; out-of-place, it passes. The severity varies between runs — 2 of 3 one time, 3 of 3 the next — which is the signature of a buffer-aliasing race rather than a logic error.
VLLM_XPU_INPLACE_ALLREDUCE is read by Intel’s stock image; this setting needs no patch. It costs about 10% of throughput at high concurrency — the price of a clone per collective — and it is not optional.
What our configuration patches, and what it does not
We should have said this in the first version. Both of our lanes run five locally patched vLLM files, bind-mounted read-only over the image’s copies. Some of the settings in this article only exist because of them. If you copy our compose file onto a stock intel/llm-scaler-vllm:0.26.0-b1 image, here is what you get:
| File | What our patch does | On a stock image |
|---|---|---|
xpu_communicator.py | Also takes the out-of-place path while a graph is being captured, and skips a work.wait() that raises “wait method cannot be used for an event associated with a command graph” during capture | We have not verified graph capture without it |
multiproc_executor.py, xpu_worker.py | Gives each worker its own ONEAPI_DEVICE_SELECTOR (the isolation below). Adds VLLM_XPU_ISOLATE_DEVICES and VLLM_XPU_ISOLATE_METHOD | Those two variables do not exist; setting them does nothing |
gemma4.py | Pre-reads a per-layer scalar after weight load, so capture does not record a device-to-host sync; for E-series models, the decoder fix described below | Gemma 4 capture fails; E2B/E4B output is wrong |
triton_attn.py | A fixed-shape buffer so image attention survives capture on the Triton path | Not needed with FLASH_ATTN, which we use |
None of these patches change a model’s numbers; each one either makes capture possible or removes a host-memory cost. We gate every patched configuration for correctness against an unpatched or reference run before recording anything.
Stopping the host-RAM mirror
Out of the box, host RAM tracks VRAM roughly 1:1. Allocating 95 GiB across the dies consumes over 100 GiB of system memory, which on a 125 GiB machine caps you at --gpu-memory-util 0.64 and rules out any model that needs the full card.
The cause is not the kernel driver. It is torch-xpu creating a single SYCL context spanning every visible device. We measured an 8 GiB device allocation three ways:
| Visibility | Host RAM cost of an 8 GiB allocation |
|---|---|
| all devices visible | +8,283 MiB |
ZE_AFFINITY_MASK=0 | +142 MiB |
ONEAPI_DEVICE_SELECTOR=level_zero:0 | −2 MiB |
The distinction between the last two matters enormously. ZE_AFFINITY_MASK filters at the Level Zero layer, and oneCCL discovers its topology through Level Zero directly — so hiding devices there leaves each rank alone in its own communication “plane” and collectives deadlock as soon as two requests are in flight. We tested nine variants; all hang.
ONEAPI_DEVICE_SELECTOR filters at the SYCL layer instead. Torch sees one device, so it builds a single-device context and no mirror; raw Level Zero enumeration stays intact, so oneCCL keeps a healthy topology. Host RAM for the dense 31B model drops from 87 GiB to 25 GiB, and --gpu-memory-util 0.90 becomes reachable.
The selector has to be set per worker — each tensor-parallel rank needs a different device — so it cannot simply go in the compose environment. That is what our multiproc_executor.py and xpu_worker.py patches do.
One caveat that cost us an afternoon. ZE_AFFINITY_MASK renumbers the devices it exposes. With ZE_AFFINITY_MASK=4,5,6,7 the four dies appear to Level Zero as indices 0–3, so the selector must name the position within the mask, not the absolute device id. A lane on dies 0–3 works either way, which hides the bug until you start a second lane somewhere else.
Graph capture, and how it ate the hypervisor
With the all-reduce fixed, graph capture becomes correct — and it is worth 3.9× on decode. It also needs one more setting, learned the hard way.
Capture’s default shape list is roughly 67 shapes. On a 120B mixture-of-experts model each captured shape pins references to the expert kernels, and that working set lives in host memory — the one resource device isolation does nothing to protect. Our first attempt consumed the machine: SSH and the Proxmox web interface both stopped being able to fork, and it needed a power cycle from the BMC. The kernel was alive the whole time; nothing in userspace could allocate.
-cc.cudagraph_mode=PIECEWISE
-cc.cudagraph_capture_sizes=[1,2,4,8,16,32,64] # 64 only if --max-num-seqs is 64Seven shapes instead of 67. Capture memory is under 2 GiB and host RAM rises about 4 GiB over eager. Adding the 64 shape cost about 1% of the KV pool and nothing measurable in host memory. Batches between captured sizes are padded up to the next one, so a batch of 50 runs in the 64-wide graph — which shows up in the NVIDIA comparison below.
bfloat16, not float16
float16 gives roughly 55% more prefill throughput on text, and we do not recommend it.
Gemma 4’s vision tower is excluded from fp8 quantisation — the recipe carries ignore: ["re:.*vision.*"] — so it inherits whatever global dtype you set. In float16 it produces unrelated descriptions of images and writes NaN into the KV cache, which then poisons later requests on the same engine. Text keeps working perfectly, which is exactly how a text-only benchmark blesses a broken configuration.
We now gate every configuration with a four-quadrant colour grid, checking the colours and their order — a model that has lost spatial understanding will still name plausible colours. That gate has caught three production-breaking configurations that text tests waved through.
Mixture-of-experts: it depends on your checkpoint
The common wisdom is that MoE models fall back to generic Triton kernels on Intel GPUs. That is true for some checkpoints and false for others, and the difference is large.
| Checkpoint quantisation | Kernel selected | Share of bandwidth roofline |
|---|---|---|
| per-channel fp8 | Triton fallback | ~20% |
| block-128 fp8 | XPUExpertsBlockFp8 | ~40% |
| MXFP4 | XPUExpertsMxFp4 | ~18% |
A native Intel MoE kernel runs at up to twice the efficiency of the Triton fallback. gpt-oss-120B — 117 billion parameters, 5.1B active, MXFP4 — is the fastest model on this box by a wide margin, and the only one that reaches a native kernel without requantisation. Its kernel is also the least efficient of the native ones: at about 18% of the memory-bandwidth roofline, there is roughly 5× of headroom left that only a better kernel from Intel can unlock. Nothing in the configuration file moves it.
Two practical notes. Block-128 quantisation constrains sharding: each shard’s dimension must still divide by 128 after the split, so a model with a 768-wide MoE intermediate dimension cannot load at TP=4 (768 ÷ 4 = 192) but loads fine at TP=2. And vLLM ships 327 tuned MoE configs for AMD and NVIDIA and none for Intel, so anything on the Triton path runs on generic defaults and says so in the log.
We also tried expert parallelism on gpt-oss (--enable-expert-parallel), which places whole experts on each die so only routed tokens cross between them. On this build it is slower everywhere: 6% on realistic chat, 31% at 32 concurrent on the standard benchmark, 12% on prefill. It does free 15% more KV cache, which is not worth that.
Performance
Aggregate decode throughput against concurrency
Prompt processing by prompt length
| gpt-oss-120B | Gemma 4 31B + vision | |
|---|---|---|
| dies / parallelism | 4–7, TP=4 | 0–3, TP=4 |
| context | 131,072 | 262,144 |
| KV cache | 876,767 tokens (fp8) | 687,047 tokens (fp8) |
| max concurrent sequences | 64 | 32 |
| decode, 1 stream | 81–87 | 32.79 |
| decode @32 | 868 (27.1/req) | 453.18 (14.2/req) |
| decode @64 | 1,116 (17.4/req) | — |
| realistic chat @32 / @64 | 1,043 / 1,444 | — |
| prefill, 8k / 30k prompt | 4,943 / 4,384 | 1,623 / 1,398 |
| time to first token @1 / @32 / @64 | 124 ms / 1,236 ms / 1,987 ms | 296 ms / 3,675 ms / — |
| vision | no vision tower | verified |
| tool calling | yes (one call at a time) | yes, including parallel |
“Decode” rows are vllm bench serve with random 512-token prompts and 256-token outputs. “Realistic chat” is our own harness: real questions, streamed, up to 400-token answers. Real chat prompts are short, so it reads higher.
Raise the cap to 64
Our first version told you the box saturates around 32 concurrent requests and to cap --max-num-seqs there. That was wrong, and the mistake is instructive: the plateau we measured was the cap itself. With the cap at 32, requests 33 and up simply queue, and throughput looks flat.
| gpt-oss-120B, 4 dies | cap 32 | cap 64 | change |
|---|---|---|---|
| realistic chat, 64 users | ~1,048 (queued) | 1,444 | +38% |
| random 512/256, 64 users | 869 (queued) | 1,116 | +28% |
| everything at 32 users or fewer | — | same | ±1% |
Raise it only if your KV cache can hold the extra sequences — with fp8 KV, gpt-oss has 876,767 tokens, which comfortably covers 64 typical chats. Past 64 the box itself saturates: a single-die Gemma 4 E4B lane, for example, is flat from 64 to 96.
fp8 KV cache
Quantising the KV cache doubles the pool for a few percent of throughput, and leaves the vision gate passing:
| Configuration | KV tokens | Full-context sessions | Cost |
|---|---|---|---|
| Gemma 31B, 128k, bf16 KV | 245,447 | 1.87× | — |
| Gemma 31B, 256k, fp8 KV | 687,047 | 2.62× | −3% @16 |
| gpt-oss, 128k, bf16 KV | 443,213 | 3.38× | — |
| gpt-oss, 128k, fp8 KV | 886,483 | 6.76× | −1% chat, −3% prefill @8k, −9% @30k |
For agent work the pool matters more than it looks. An agent framework with a 47,000-token system prompt and tool list fits about nine concurrent sessions in gpt-oss’s bf16 pool and eighteen in fp8.
Both lanes at once
Running the two four-die lanes simultaneously is essentially free. Under 16-concurrent load on each over the same window, one lane retained 102% of its solo throughput and the other 101% — the excess is run-to-run variance; the honest reading is no measurable contention. Combined output was 1,297 tokens per second on our original configuration. The same holds for single-die lanes: two of them at 32 concurrent each showed no decode slowdown, though first-token time rose on the slower one as they competed for host CPU.
That was not a given. The lanes share root complexes, host memory bandwidth, the PCIe fabric and 449 GiB of committed MMIO on a CPU with 43 physical address bits. A 20–30% mutual penalty would not have surprised us.
Small Gemma 4 models, one die each — and a bug in vLLM
Gemma 4 also comes in small sizes, and a small model on a single die avoids tensor-parallelism entirely: no all-reduce, no device isolation, no capture aliasing. We tested two, the E4B (about 4B effective parameters, fp8) and Google’s 12B quantisation-aware INT4 release.
The E4B looked spectacular at first — over 1,000 tokens per second on one die — and passed our three-question correctness gate. Then its image answers came back as “an, green, blue, you”, and a closer look showed every long answer collapsing into loops: “kind of kind of kind of…”. Short gate answers had hidden it. Google’s own bf16 weights did the same, so the quantisation was not to blame, and none of graph capture, KV dtype, attention backend, prefix caching, block size, sampling or Transformers version changed anything.
Hugging Face’s own implementation, on the same PyTorch and the same GPU, wrote clean answers. Feeding its answers back through vLLM and comparing token by token, vLLM disagreed from the first token — average log-probability error of 2 to 4 — which pointed at something computed for every token. It was this, in the decoder layer, only on the E-series’ per-layer-embedding branch:
# Hugging Face Gemma4TextDecoderLayer (reference)
hidden = residual + post_ff_norm(mlp_out) # add the residual first
gate = per_layer_input_gate(hidden) # gate sees the full stream
hidden = hidden + post_ple_norm(ple_proj(gelu(gate) * per_layer_input))
# vLLM 0.26 Gemma4DecoderLayer, as shipped
hidden = post_ff_norm(mlp_out) # residual NOT yet added
gate = per_layer_input_gate(hidden) # gate sees only the FFN delta
hidden = hidden + post_ple_norm(ple_proj(gelu(gate) * per_layer_input))
hidden = (hidden + residual) * layer_scalarThe gate in every one of E4B’s 42 layers was reading the wrong tensor. Moving the residual add above the gate fixes it: vLLM then matches Hugging Face to within rounding (average error 0.01–0.1), long answers are clean, tools work including parallel calls, and the colour grid comes back “red, green, blue, yellow”. Models without per-layer embeddings — 12B, 26B, 31B — never take that branch, which is why only the E-series broke. This is an upstream vLLM bug, not an Intel one.
| Realistic chat, one die | Gemma 4 E4B fp8 (fixed) | Gemma 4 12B QAT INT4 |
|---|---|---|
| KV cache, 128k context | 847,015 tokens | 434,376 tokens |
| 1 user | 60.7 tok/s | 32.8 tok/s |
| 16 users, total (per user) | 844 (54.0) | 395 (25.2) |
| 32 users | 1,433 (46.0) | 738 (23.8) |
| 64 users (cap 64) | 2,370 (38.4) | — |
| prefill 8k / 30k | 6,511 / 4,350 | 2,053 / 1,551 |
| vision, tools | yes, yes | yes, yes |
One die of E4B out-serves the whole four-die 31B lane roughly threefold, and one die of the 12B matches it. The 12B is the better writer; the E4B is the better throughput part.
Real agent prompts change the picture. These figures use short chat prompts. A real agent workload against the same E4B die, with roughly 5,000-token prompts and a prefix-cache hit rate of only 10–18%, measured 675 tokens per second at 32 users — 21 per user. Once prompts are long, reading them dominates, and prefix caching only helps if the start of every prompt is byte-identical. A timestamp or user name near the top of a system prompt defeats it.
The 12B needs a specific Transformers version
Every Gemma 4 12B checkpoint, including the community quantisations, uses the config type gemma4_unified. The Transformers 5.8 in Intel’s image does not know it and refuses to start. Transformers 5.15 and later know it, but describe head_dim per layer (256 on sliding-window layers, 512 on full-attention ones) and drop the global field this vLLM reads; forcing access would build the full-attention layers wrong. Transformers 5.10 to 5.14 know the model and keep the old field. We mount 5.14.1 into the 12B lane only, prepended with PYTHONPATH; it runs on the image’s own tokenizers, and every other lane keeps 5.8.
Against NVIDIA, honestly
Aggregate memory bandwidth comparisons are architecturally interesting and tell you very little about serving. The comparison that matters is same model, same memory, published numbers — and the natural opponent is the card with identical VRAM.
96 GB against 96 GB
Two MAXSUN dual-die B60 cards give four GPUs and 96 GB of VRAM. One NVIDIA RTX PRO 6000 Blackwell gives one GPU and 96 GB of VRAM. Same capacity, very different invoice.
Rather than compare our sweep against a benchmark run under different conditions, we reproduced the published RTX PRO 6000 harness2 — same model, same framework, 100-token prompts, 600-token generations, every request forced to the full length, 50 concurrent requests — and re-measured on our tuned lane.
| 2 × MAXSUN Dual B60 | 1 × RTX PRO 6000 Blackwell | |
|---|---|---|
| VRAM | 96 GB (4 × 24) | 96 GB |
| parallelism | TP = 4 | TP = 1 |
| quantisation | MXFP4 | 8-bit |
| price, AUD | $5,398 | $25,100 |
| output tok/s @ 50 users | 1,094 | 1,779 |
| per user | 21.9 tok/s | 35.6 tok/s |
| share of NVIDIA throughput | 61% | — |
| tok/s per AUD 1,000 | 203 | 71 |
| output tok/s @ 64 users | 1,510 | not published |
gpt-oss-120B — same harness, throughput per thousand dollars spent
What is and is not matched. Same model, same inference engine, same prompt and generation lengths, same concurrency, all requests completing. Three differences remain. Their card runs the model on one GPU where ours needs four, which is the comparison being made rather than a flaw in it. Their quantisation is 8-bit against our MXFP4, each being the vendor-native path for this model on that hardware. And this re-run used our production 131,072-token context, not their 4,096 — stricter on us, but with 876,767 KV tokens the 35,000 this test needs exert no pressure either way. The 50-user figure is the mean of two warm runs (1,085 and 1,103); a cold first run read 1,007 and is excluded.
Two things the re-run showed. At 32 concurrent the lane now does 1,149 tok/s, up from 1,077. And 64 users get 1,510 — more than 50 — because a batch of 50 is padded into the 64-wide captured graph and pays for 64. NVIDIA has not published a 64-user figure, so we do not compare it; we show it because it is the number this hardware actually delivers when you size the cap for it. A captured shape at 48 might lift the 50-user figure; we have not tried it.
Against a datacentre card
SemiAnalysis’s InferenceX publishes per-chip gpt-oss-120B figures at 8K context and 1K generation in FP4,1 which brackets our numbers usefully.
| Configuration | Per-user tok/s | Concurrent users | Precision |
|---|---|---|---|
| 1 × H100 (SemiAnalysis) | 197 | ~6 | FP4 |
| 1 × H100 (SemiAnalysis) | 143 | ~13 | FP4 |
| 1 × H100 (SemiAnalysis) | 90 | ~32 | FP4 |
| 4 × Arc Pro B60 (measured here) | 85 | 1 | MXFP4 |
| 4 × Arc Pro B60 (measured here) | 39.2 | 16 | MXFP4 |
| 4 × Arc Pro B60 (measured here) | 27.1 | 32 | MXFP4 |
| 4 × Arc Pro B60 (measured here) | 17.4 | 64 | MXFP4 |
At 32 concurrent users a single H100 sustains about 90 tokens per second per user; four B60 dies sustain 27.1 under our profile. That is roughly 30% of one H100’s per-user throughput at the same concurrency. At a single stream the four dies match an H100’s 32-user operating point — low concurrency is where this hardware flatters itself. This comparison is less controlled than the RTX PRO 6000 one above, since we did not replicate the SemiAnalysis harness.
The case for these cards is economic, not technical. If you are buying throughput per dollar for fixed internal serving, roughly a third to two thirds of the performance for a fifth of the price is a good trade — and 191 GiB of VRAM in one chassis lets you run models that will not fit on a single accelerator at any speed. If you are buying per-user latency at high concurrency, buy the NVIDIA card. Nothing here makes this a datacentre-class part and we would not pitch it as one.
Two caveats. SemiAnalysis measure NVIDIA FP4 against our MXFP4 — both four-bit, not identical. Their harness uses 8K context and 1K generation; our sweep uses a different prompt profile. Treat it as indicative of the band rather than a controlled head-to-head.
The whole box
Scaled to the full chassis, the arithmetic is the reason we built it this way:
| 4 × MAXSUN Dual B60 | 2 × RTX PRO 6000 | |
|---|---|---|
| VRAM | 191 GiB | 192 GB |
| GPUs | 8 | 2 |
| price, AUD | $10,796 | $50,200 |
| configuration | 2 independent TP=4 lanes | 1 TP=2 lane, or 2 × TP=1 |
| gpt-oss-120B, matched harness | 1,094 tok/s per 4 dies | 1,779 tok/s per card |
Roughly AUD 39,000 of difference for the same memory. That buys a great deal of engineering time to work around an immature software stack — which, as the rest of this article documents, is exactly what it cost us. Whether that trade makes sense depends entirely on whether your engineering time is cheaper than AUD 39,000, and on whether you need the throughput you are giving up.
The configuration, in full
This is the production compose file for the vision lane, as it runs today. The /opt/vllm/patched mounts are the local patches described above.
services:
lane-a:
image: intel/llm-scaler-vllm:0.26.0-b1
container_name: lane-a
network_mode: host
ipc: host
privileged: true
init: true
restart: "no" # a crash-loop hides its own failure; an external
# watchdog restarts a dead or hung lane instead
shm_size: 32g # the 64MB default causes collective hangs
devices:
- /dev/dri:/dev/dri
ulimits:
# oneCCL pins memory for its fabrics; the default memlock is far too low
memlock: {soft: -1, hard: -1}
stack: {soft: 67108864, hard: 67108864}
volumes:
- /models:/models:ro
# vLLM ships 327 tuned MoE configs for AMD and NVIDIA and zero for Intel.
# fused_moe.py checks VLLM_TUNED_CONFIG_FOLDER first, so a generated JSON
# here overrides the shipped set without rebuilding the image.
- /opt/vllm/moe-tuned:/moe-tuned
# local patches - see "What our configuration patches"
- /opt/vllm/patched/xpu_communicator.py:/opt/venv/lib/python3.12/site-packages/vllm/distributed/device_communicators/xpu_communicator.py:ro
- /opt/vllm/patched/gemma4.py:/opt/venv/lib/python3.12/site-packages/vllm/model_executor/models/gemma4.py:ro
- /opt/vllm/patched/xpu_worker.py:/opt/venv/lib/python3.12/site-packages/vllm/v1/worker/xpu_worker.py:ro
- /opt/vllm/patched/multiproc_executor.py:/opt/venv/lib/python3.12/site-packages/vllm/v1/executor/multiproc_executor.py:ro
- /opt/vllm/patched/triton_attn.py:/opt/venv/lib/python3.12/site-packages/vllm/v1/attention/backends/triton_attn.py:ro
environment:
- VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- CCL_TOPO_P2P_ACCESS=1
# Default is 300s. A 64k prompt on the Triton fallback exceeds it and the
# engine dies (EngineDeadError) - one long request takes down the server
# for everyone. NB: VLLM_RPC_TIMEOUT is NOT a real variable in this image;
# setting it fails silently and looks exactly like a fix that did not work.
- VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=3600
# Per-worker SYCL-layer isolation. These two variables come from OUR
# patched executor/worker; a stock image ignores them.
- VLLM_XPU_ISOLATE_DEVICES=1
- VLLM_XPU_ISOLATE_METHOD=selector
- VLLM_XPU_ENABLE_XPU_GRAPH=1
# MANDATORY. Without this, TP=4 graph capture silently corrupts output.
- VLLM_XPU_INPLACE_ALLREDUCE=0
- VLLM_TUNED_CONFIG_FOLDER=/moe-tuned
- ZE_AFFINITY_MASK=0,1,2,3
command:
- '/models/gemma-4-31b-fp8-dyn'
- '--tensor-parallel-size'
- '4'
- '--dtype'
- 'bfloat16' # NOT float16 - fp16 destroys the vision tower
- '--block-size'
- '64'
- '--gpu-memory-util'
- '0.90'
- '--max-model-len'
- '262144'
- '--max-num-batched-tokens'
- '8192'
- '--enable-prefix-caching'
- '--kv-cache-dtype'
- 'fp8_e4m3' # doubles the KV pool for a few percent
- '-ac.backend=FLASH_ATTN'
- '--max-num-seqs'
- '32'
# Six capture shapes, not the default ~67. The default list wedged the
# hypervisor on a 120B MoE badly enough to need a power cycle.
- '-cc.cudagraph_mode=PIECEWISE'
- '-cc.cudagraph_capture_sizes=[1,2,4,8,16,32]'
- '--enable-auto-tool-choice'
- '--tool-call-parser'
- 'gemma4'
- '--reasoning-parser'
- 'gemma4' # without this, channel markup leaks into content
- '--host'
- '0.0.0.0'
- '--port'
- '8000'The gpt-oss lane differs only here
- ZE_AFFINITY_MASK=4,5,6,7
- '/models/gpt-oss-120b'
- '--max-model-len'
- '131072'
- '--kv-cache-dtype'
- 'fp8_e4m3' # 443,213 -> 886,483 tokens for ~1% of chat throughput
- '--max-num-seqs'
- '64' # not 32 - see "Raise the cap to 64"
- '-cc.cudagraph_capture_sizes=[1,2,4,8,16,32,64]'
- '--tool-call-parser'
- 'openai' # NOT "gptoss" - filename and registered name differ
# no --reasoning-parser: gpt-oss's is selected automatically
- '--port'
- '8001'What does not work
| Configuration | Result |
|---|---|
--dtype float16 | Destroys the vision tower. Text unaffected, which is how it passes a text-only benchmark. |
Graph capture without VLLM_XPU_INPLACE_ALLREDUCE=0 | Fluent, confident, wrong output at full speed. |
Default cudagraph_capture_sizes on a 120B MoE | Exhausts host memory; needed a power cycle. |
ZE_AFFINITY_MASK for device isolation | Removes the mirror but gives one rank per oneCCL plane; deadlocks at concurrency 2. |
--enable-expert-parallel on gpt-oss | 6% slower on chat, 31% slower at 32 concurrent, 12% slower prefill. |
| Gemma 4 E2B/E4B on stock vLLM 0.26 | Loops on any answer longer than a sentence; blind to images. Needs the decoder fix above. |
| Gemma 4 12B with Transformers 5.8, or 5.15 and later | Refuses to start, or builds full-attention layers wrong. Use 5.10–5.14. |
| Pipeline parallelism (4×1, 1×4) | Will not start. 2×2 costs 58% throughput. |
--block-size 128 or 256 | No measurable effect. |
--gpu-memory-util 0.95 | Worker init fails. 0.90 is the practical ceiling. |
CCL_ZE_IPC_EXCHANGE=drmfd | Kills the engine. |
| AWQ 4-bit weights on the 31B | Slower than fp8 despite 40% fewer reads. |
| BIOS tweaks (ACS, C-states, TSME) | Performance-neutral. |
| Per-channel fp8 MoE checkpoints | Forced onto the Triton fallback; roughly half the efficiency of a native kernel. |
Environment variables that do not exist in this build, and fail silently if you set them: VLLM_ATTENTION_BACKEND, VLLM_XPU_ENABLE_FP8, VLLM_USE_V1, VLLM_RPC_TIMEOUT — and, without our patches, VLLM_XPU_ISOLATE_DEVICES and VLLM_XPU_ISOLATE_METHOD. Setting a variable vLLM does not read looks exactly like a fix that did not work, which cost us more time than we would like to admit.
Five traps when you evaluate these models
None of these is about the hardware, but each cost us real time and each would silently corrupt anyone else’s comparison.
The models default differently on reasoning
gpt-oss always reasons — its reasoning_effort parameter is inert in this build, with low, medium and high returning byte-identical output. Gemma 4’s chat template defaults enable_thinking to false, so it does not reason unless you ask.
Our first capability comparison therefore had one model thinking and the other not, and concluded that Gemma was poor at multi-step arithmetic. It was not. It had never been asked to think:
gemma, thinking off (default): wrong 6 tok 0.2s "55.11"
gemma, thinking on: CORRECT 783 tok 23.8s "51.74"
# enable it per request:
"chat_template_kwargs": {"enable_thinking": true}Corrected, the two models score identically on every mechanically checked case. But thinking is a trade rather than an upgrade — on a strict-format task (“exactly three sentences, no letter z, end with DONE”) Gemma with thinking enabled spent 60 seconds and its entire token budget spelling out every word letter by letter to verify the constraint, and never produced an answer. With thinking off it answered correctly in 1.1 seconds.
Check template defaults, not just which parsers are registered. We verified the reasoning and tool parsers carefully and never looked at what the template does when the flag is absent. That is where the difference was.
gpt-oss rejects “reasoning off”
Its Harmony format accepts a reasoning_effort of low, medium or high, and nothing else. Sending "none" — which many agent frameworks do for background calls such as conversation titles, to keep them quick — returns HTTP 400: Harmony does not support reasoning_effort=’none’. The main conversation works and the background task fails quietly. We route those calls to the Gemma lane, which accepts it.
gpt-oss emits non-ASCII characters inside shell commands
This one no automated score would ever catch, and we only found it by reading the open-ended answers rather than grading them. Asked to produce diagnostic commands, gpt-oss returns:
Test‑NetConnection -ComputerName mail.example.com -Port 443
^
U+2011 NON-BREAKING HYPHEN, not an ASCII hyphenFifteen non-breaking hyphens and twenty-eight smart quotes in a single answer. The command looks correct in any rendered view and fails on execution. Gemma emits clean ASCII in commands in both modes.
If you feed gpt-oss output to anything that executes it, normalise first. An agent that runs generated commands will fail on text that looks perfectly valid to a human reviewing the transcript. unicodedata.normalize("NFKC", text) and a hyphen/quote fold before execution.
Short correctness gates miss broken models
The broken E4B answered “Paris”, “4” and “blue” perfectly. Its failure only appeared after 30 to 700 tokens. Every gate now includes a 1,024-token answer checked for repetition, and a token-by-token comparison against a reference implementation whenever a model is new to the stack.
Random-token benchmarks cannot measure speculative decoding
Google publishes small draft models that let a Gemma 4 model guess several tokens ahead and keep the ones it would have produced anyway. On the 12B, vllm bench serve with random prompts measured the draft as a heavy loss — because random text cannot be guessed, its acceptance rate was 3%. On real questions it was 55%:
| 12B QAT, one die, real prompts | without draft | with draft (2 ahead) | change |
|---|---|---|---|
| 1 user | 32.8 tok/s | 53.3 tok/s | +62% |
| 4 users | 127 | 135 | +6% |
| 16 users | 395 | 287 | −27% |
| 32 users | 738 | 345 | −53% |
Both halves matter. Measure speculative decoding on real text, and use it only where few users share a die: every rejected guess is work the whole batch pays for.
What we would tell you
The hardware is capable and the software stack is immature — but nearly every limit we originally attributed to the cards turned out to be a software defect with a fix. Prompt processing was not bandwidth-bound, it was kernel-bound. Multimodal was not broken, it was mis-reduced. Mixture-of-experts was not unsupported, it was a quantisation-format mismatch. The throughput ceiling was a setting we had chosen. And the small Gemma models were not weak, they were mis-wired — upstream, on every platform.
If you deploy on these: use bfloat16, set VLLM_XPU_INPLACE_ALLREDUCE=0, cap your capture shapes, isolate devices at the SYCL layer, size --max-num-seqs to your KV pool rather than to a rule of thumb, and above all gate every configuration on correctness — with long answers — before you record its throughput. The most dangerous failure on this stack is not a crash. It is a model that sounds right.
References
- SemiAnalysis InferenceX, gpt-oss 120B — B200 vs H100 Inference Benchmark. Per-chip figures at 8K context / 1K generation, FP4 precision.
- Database Mart, RTX PRO 6000 vLLM Inference Benchmark. GPT-OSS-120B at 1,779 tok/s aggregate, single GPU, tensor-parallel-size 1, 8-bit, 50 concurrent requests, 100-token prompts, 600-token generations,
--max-model-len 4096.
Hardware prices are Australian retail as at September 2026: MAXSUN Arc Pro B60 Dual 48 GB at AUD 2,699, NVIDIA RTX PRO 6000 Blackwell 96 GB at AUD 25,100. Prices move; the ratio is the point.
All figures were measured on the hardware named in the byline and gated for correctness before being recorded. Configurations that failed a gate were discarded rather than reported, including several that were faster. First published 17 September 2026; revised 21 September and 1 October 2026.