THE RESULT One NVIDIA DGX Spark kept Qwen 3.6 27B and Gemma 4 31B resident behind a shared Minima endpoint. Two teams received separate Kubernetes clusters through vCluster, and a Job in each cluster called its assigned model. After team-qwen was deleted and rebuilt, team-gemma remained Running and the Minima router still answered requests.
A compact GPU system is easier to share when teams can use different models without waiting for a model swap and can deploy their own workloads without coordinating every Kubernetes change. Minima and vCluster tested those two requirements together on a Spark that was already serving models.
Minima operated the model service on the host. vCluster Standalone provided the Kubernetes control plane cluster, with the Spark joined as its worker. From there, the operator created one tenant cluster for team-qwen and another for team-gemma. Each team submitted a small Job through its own Kubernetes API; both Jobs reached the same Minima router over the Pod network.
The source experiment also measured the model endpoint under serial, concurrent, and mixed-model requests. Those measurements describe the configured Minima service. The tenant Jobs separately demonstrate access from the two clusters. There is no matched stock-vLLM performance baseline in this run.
The System We Built
The only GPU system in the test was an arm64 NVIDIA DGX Spark with a GB10 processor, running Ubuntu 24.04. Docker, containerd, Tailscale, and Minima’s model services were present before vCluster was installed.
Minima’s host-managed router listened on port 8000. It selected a resident backend from the model field of an OpenAI-compatible request:
| Requested model ID | Backend port | Reported max_model_len |
|---|---|---|
qwen3.6-27b | 8867 | 32,768 |
gemma4-31b-it | 8868 | 32,768 |
A client could therefore use one /v1 address for either model. The two backends remained loaded, so changing the model field did not require unloading one to make room for the other.
vCluster Standalone runs a control plane binary directly on a machine without requiring an underlying control plane cluster. Worker nodes join it as private nodes. In this test, the Standalone instance on the Spark also served as the control plane cluster for two further tenant clusters. team-qwen and team-gemma each received its own API server, CoreDNS, kubeconfig, and Kubernetes view.
Each team’s workload was a Kubernetes Job using a hardened curl container. The Job discovered the node address through status.hostIP and posted one deterministic prompt to the model assigned to that team. It represented the kind of client a team might otherwise deploy as an application, agent, or batch pipeline.
The separation has a precise limit. The teams had distinct control planes and Kubernetes views, but their workloads shared one machine, one GPU, and one kernel. This experiment did not give each tenant its own worker node or establish hard isolation between external customers.
Why Minima Fits This Hardware
Minima optimizes model weights, the KV cache, hardware-specific operations, and the serving runtime. Its model-compression pipeline identifies compressible structure, applies tensor decompositions, repairs quality with a short fine-tune, and serves the resulting artifact through an optimized runtime.
The Spark’s Blackwell-based GB10 processor made it a useful target for this work: the experiment asked whether two capable models could remain available within one compact system’s resource envelope. The result was the two resident backends used throughout the run.
A separate, previously published Minima evaluation provides context, but it is not a measurement from this Spark experiment. For Qwen3-32B at 8K context, Minima reported reducing peak VRAM from 64 GiB to 40 GiB while increasing throughput in a controlled A100 setup.
Verifying the Shared Endpoint
The operator first checked the model service directly on the Spark:
export MINIMA_BASE_URL="http://127.0.0.1:8000/v1"
curl -fsS "$MINIMA_BASE_URL/models" | python3 -m json.tool
The returned model list included gemma4-31b-it and qwen3.6-27b, each with max_model_len=32768. Here 127.0.0.1 names the Spark itself. The same address would name a different machine if used on a laptop or inside a Pod.
A direct /chat/completions request selected Qwen with "model": "qwen3.6-27b". It used a fixed prompt requesting Qwen on Spark is ready, temperature: 0, max_tokens: 32, and chat_template_kwargs: {"enable_thinking": false}. The router returned the requested text. Changing the model ID and prompt in the same request shape returned Gemma on Spark is ready.
Gemma also received the architecture diagram in a separate vision check. It described the router, both backends and ports, and the model-field routing.
The shell export mattered operationally: MINIMA_BASE_URL applied only to that shell session and had to be set again in a new session or after switching users with sudo -i.
What the Model Endpoint Delivered
Before Kubernetes was added, the team measured the service that the tenant workloads would later call. The client opened streaming requests against the router, timed the first content event and subsequent decode rate, and recorded raw per-request JSON. It did not restart or reconfigure either backend during the measurements.
The serial profile used one warm-up followed by five measured runs per model, with 128 generated tokens:
| Model | Median time to first token | Median decode rate | Median request time |
|---|---|---|---|
| Qwen 3.6 27B | 186 ms | 25.96 tokens/s | 5.07 s |
| Gemma 4 31B | 296 ms | 18.84 tokens/s | 7.02 s |
The same script then ran three batches of 16 simultaneous 128-token requests per model:
| Model | Completed requests | Median aggregate output rate at concurrency 16 |
|---|---|---|
| Qwen 3.6 27B | 48/48 | 231.06 tokens/s |
| Gemma 4 31B | 48/48 | 76.21 tokens/s |
A third profile called Qwen and Gemma at the same time in five paired batches. Every pair completed. Median aggregate output rate for that mixed-model profile was 19.20 tokens/s; Qwen decoded at roughly 11.1 tokens/s and Gemma at roughly 10.0 tokens/s. The mixed-model rate belongs to its own request pattern and should not be read as the sum of the separate concurrency-16 results.
The script’s captured Qwen serial summary gives more detail about the five measured requests:
| Qwen serial metric | Median | p95 | Minimum | Maximum |
|---|---|---|---|---|
| Time to first token | 0.186 s | 0.192 s | 0.152 s | 0.194 s |
| Decode rate | 25.96 tokens/s | 26.05 tokens/s | 25.79 tokens/s | 26.07 tokens/s |
| Total request time | 5.072 s | 5.104 s | 5.050 s | 5.111 s |
These are client-observed results for the service as Minima handed it over, rather than figures reported by the runtime itself. The benchmark script ran from a laptop using the Spark’s Tailscale address:
export MINIMA_BASE_URL="http://<SPARK_TAILSCALE_IP>:8000/v1"
python3 scripts/minima_stream_benchmark.py \
--base-url "$MINIMA_BASE_URL" \
--max-tokens 128 \
--serial-runs 5 \
--concurrency 16 \
--concurrent-batches 3
Its JSON output contains the warm-up, each request, and per-model summaries. The experiment did not include a matched stock-vLLM arm using the same artifacts, prompts, cache state, and client. It therefore supports no measured speedup claim over stock vLLM. Nor do these laptop-to-router results measure throughput through the tenant clusters.
Installing vCluster Beside Running Models
The Spark was not an empty test host. Before installing vCluster, the operator recorded what needed to keep working: both models had to remain listed under /v1/models; the router had to stay reachable locally and over LAN; SSH had to work over LAN and Tailscale; and the existing Docker and containerd runtime had to remain healthy.
The run pinned vCluster v0.36.0 and Kubernetes v1.36.0. Its Standalone configuration addressed three constraints of the occupied host: use the existing containerd CRI socket, retain host swap for the memory-heavy model processes while preventing Pods from using it, and join the control-plane host as the single worker.
controlPlane:
standalone:
advertiseAddress: <SPARK_LAN_IP>
joinNode:
enabled: true
containerd:
enabled: false
preJoinCommands:
- systemctl start containerd
nodeRegistration:
criSocket: unix:///run/containerd/containerd.sock
distro:
k8s:
image:
tag: v1.36.0
privateNodes:
enabled: true
kubelet:
config:
failSwapOn: false
memorySwap:
swapBehavior: NoSwap
The reviewed configuration was copied to /etc/vcluster/vcluster.yaml on the Spark. Installation used the pinned Standalone release and named the instance spark-ai:
curl -sfL \
https://github.com/loft-sh/vcluster/releases/download/v0.36.0/install-standalone.sh \
| sudo sh -s -- \
--vcluster-name spark-ai \
--config /etc/vcluster/vcluster.yaml
The Pod network also needed a path to services on the host’s private LAN address. With 10.244.0.0/16 as the Pod CIDR, the run allowed traffic arriving on cni0 to TCP port 8443 for the Standalone API, 8000 for Minima’s router, and 10250 for the kubelet:
export SPARK_LAN_IP="<SPARK_LAN_IP>"
export POD_CIDR="10.244.0.0/16"
sudo ufw allow in on cni0 from "$POD_CIDR" to "$SPARK_LAN_IP" \
port 8443 proto tcp comment 'vcluster pods to standalone api'
sudo ufw allow in on cni0 from "$POD_CIDR" to "$SPARK_LAN_IP" \
port 8000 proto tcp comment 'vcluster pods to minima router'
sudo ufw allow in on cni0 from "$POD_CIDR" to "$SPARK_LAN_IP" \
port 10250 proto tcp comment 'vcluster agent to kubelet'
Those values describe the tested setup; the private LAN address is intentionally a placeholder. Although the Standalone configuration enables privateNodes, this single-Spark run did not assign a separate physical node to each of the two team clusters.
One Worker, Two Kubernetes Views
With KUBECONFIG=/var/lib/vcluster/kubeconfig.yaml, kubectl get nodes -o wide showed the Spark Ready as the control-plane and worker node. The captured node output was taken a week after installation and reported Kubernetes v1.36.0 with containerd://2.2.1. Flannel, CoreDNS, the Konnectivity agent, kube-proxy, and the local-path provisioner were Running. The local-path provisioner mattered because each tenant control plane claimed a volume.
The operator then created the two clusters:
vcluster create team-qwen --namespace team-qwen --add=false
vcluster create team-gemma --namespace team-gemma --add=false
--add=false skipped registration with vCluster Platform, which was not installed. The captured status showed team-gemma Running at 52 seconds and team-qwen Running at 55 seconds, both on vCluster 0.36.0.
Inside team-qwen, a namespace listing contained default, kube-node-lease, kube-public, and kube-system. It did not expose team-gemma or the underlying cluster’s Flannel and local-path namespaces. team-gemma had the corresponding view of its own four namespaces and did not show team-qwen.
Both clusters reported a Ready node object named spark-5385, running Kubernetes v1.36.0; the captured tenant views showed addresses 10.104.23.9 for team-qwen and 10.97.29.108 for team-gemma. Those were two views of the same physical Spark, not two worker machines. The teams interacted with their tenant APIs while the operator retained visibility into the underlying control plane cluster.
Each Team Called Its Model
The workload manifests omitted a namespace field because each was applied through its team’s own cluster connection:
vcluster connect team-qwen --namespace team-qwen \
-- kubectl apply -f team-qwen-workload.yaml
vcluster connect team-gemma --namespace team-gemma \
-- kubectl apply -f team-gemma-workload.yaml
Each Job completed 1/1 in three seconds. Logs read from within the corresponding tenant cluster showed:
| Team and Job | Model | Returned text | Reported token usage |
|---|---|---|---|
team-qwen / ask-qwen | qwen3.6-27b | hello from team-qwen | 21 prompt, 6 completion, 27 total |
team-gemma / ask-gemma | gemma4-31b-it | hello from team-gemma | 23 prompt, 7 completion, 30 total |
The underlying control plane cluster showed each team’s translated workload name, its CoreDNS Pod, and its tenant control plane Pod, all scheduled on spark-5385. The operator saw the shared node and its Pods; each team worked through its own Kubernetes API.
The three-second Job completions verify that these particular Jobs finished. The endpoint benchmark above used separate streaming requests from a laptop.
Deleting a Team Without Deleting Its Neighbor
The next test removed the entire Qwen tenant cluster:
vcluster delete team-qwen --namespace team-qwen
The cluster disappeared within a few seconds. Immediately afterward, team-gemma still reported Running, and its ask-gemma Job still reported Complete. The operator recreated team-qwen, reapplied its workload manifest, and received the same hello from team-qwen answer. The source describes the delete-and-rebuild sequence as taking under half a minute.
Finally, the operator checked the original host services again. /v1/models still listed both backends with max_model_len=32768, and a live Qwen completion through the router returned Qwen on Spark is ready. SSH, Tailscale, vCluster, kubelet, containerd, and Docker were active. The pre-existing Docker model runner was reported Up for 17 hours. Host memory ended at 16 GiB available out of 121 GiB.
One team’s cluster was removed and restored while the other team’s cluster remained Running, and the host-managed model service answered at the final check.
What Was Tested, and What Comes Next
vCluster’s wider product architecture has more layers than this Spark used. vCluster can place tenant clusters on an existing Kubernetes cluster, run a control plane directly on bare metal or a VM as Standalone does, or run in Docker for local development and CI. The test exercised the Standalone control plane cluster, the two tenant clusters, and their workloads on one machine.
For a fleet, the source describes vCluster Platform as the management plane and UI across one or more control plane clusters, covering projects, templates, access, capacity policy, observability, and billing. It describes vNode as a lightweight per-Pod sandbox using Linux user namespaces, without a hypervisor or guest kernel, for a stronger boundary around customer-supplied code. vMetal offers a consistent API for bare metal servers and VMs while orchestrating provisioning, PXE booting, networking, storage, and virtualization tools.
Those are product descriptions and a proposed architecture, not components validated in this lab. The source’s production direction gives tenant clusters their own Private Nodes when the tenants are external customers, adds vNode where customer-supplied code runs, and uses Platform to manage the fleet. vMetal becomes relevant if the operator needs automated physical or virtual machine lifecycle management. It was not added to this test because the Spark already existed and had been prepared by hand.
| Tested on this Spark | Production direction described by the source |
|---|---|
| One host-managed Minima model service with Qwen and Gemma resident | A provider-operated model layer across a GPU fleet |
| One vCluster Standalone control plane cluster | vCluster Platform across control plane clusters |
| Two tenant clusters sharing one Spark worker | Each tenant cluster on its own Private Nodes |
| Jobs calling the shared model router | vNode where customer-supplied code runs; optional vMetal for machine lifecycle |
A separate Kubernetes API gives each team control of its cluster. It does not change the fact that these two teams shared a kernel and GPU. The source treats that as a boundary for the internal-team demonstration, not as hard isolation for competing customers.
How Minima Packages the Model Layer
The optimization applied to Qwen 3.6 27B and Gemma 4 31B combines structural and non-structural techniques for memory-constrained inference on GB10-class hardware. Minima quantizes both the primary model and its speculative decoding model and uses residual error compensation for quantization-induced error. Tensor-network methods reduce the model footprint; Minima KV compresses the cache; and GB10-tuned kernels handle performance-critical operations. Applying the path to both the main model and the speculator prevents an unchanged speculator from offsetting savings in the main model.
Minima says it evaluates each optimized artifact against its original, unmodified model with lm_eval: the original runs first, then the optimized artifact runs under the same evaluation configuration. Across the suite used for these models, the optimized versions reportedly stayed within approximately one percentage point of the originals.
Serving uses a custom Minima build of vLLM. The source describes its GB10-specific changes as reusable across GB10 devices rather than specific to this Spark, and the runtime as designed for a near drop-in change from stock vLLM: switch the build and add a small number of environment variables while retaining the serving workflow. That is an operational description, not a measured stock-vLLM speedup.
The described lifecycle keeps model artifacts and runtime versions together. Health checks verify loading and endpoint responsiveness; an upgrade deploys a new versioned combination; a rollback restores a known-good one; and a failed model process can restart and reload its selected artifact without repeating optimization.
For a future stock-vLLM performance comparison, the source specifies a matched baseline: the same GB10 hardware, source model, workload, context length, concurrency, request distribution, and measurement method, with the original model on stock vLLM compared against the optimized model and Minima serving path. That comparison was not included in the Spark benchmark above.
Davyd Maiboroda, Minima’s Head of Research and Founder, frames the practical goal as making constrained devices useful for capable models while addressing model swapping and multi-model serving. In this experiment, keeping Qwen and Gemma resident let different engineers, agents, or services choose between them without repeatedly unloading and reloading a backend.
Reproducing the Run
The companion repository contains the tested Standalone configuration in manifests/vcluster-standalone.yaml, one workload manifest for each team, scripts/minima_stream_benchmark.py, and demo-curls.md with sanitized text, vision, and reasoning requests.
Unless a step specifies otherwise, the installation and cluster commands run in an SSH shell on the Spark. The optional network benchmark runs from a laptop. That distinction changes what 127.0.0.1 means. The published material intentionally excludes private model paths, credentials, join tokens, and Minima implementation details that were not cleared for publication.
The source also points readers to vCluster’s inference platform guide, Standalone installation documentation, worker-node model guide, and Platform documentation for the tenancy and fleet concepts discussed here.
From One Spark to a Larger System
On the tested Spark, Minima made two model backends available without a swap, while vCluster let two teams submit workloads through separate Kubernetes control planes. The team Jobs reached their models; one cluster could be rebuilt; and the model endpoint still answered at the end. That is the demonstrated result.
The proposed extension changes the infrastructure beneath the teams: a GPU fleet, provider-operated models, Private Nodes for customer tenants, and vCluster Platform to manage the collection. vNode and vMetal have the conditional roles described above. None was exercised in the single-Spark experiment.
Minima’s source article also describes a separate frontier-model direction. A Qwen3.8 2.4T deployment is in development, with Minima targeting one 8xB200 node; the source states that the unoptimized model needs at least four 8xB200 nodes in the same serving configuration. This is a target and comparison presented by Minima, not a completed result from the Spark. In a discussion with vCluster co-founder and CEO Lukas Gentele at the AI Infra Summit, Minima co-founder and CEO Sergii Kozyrev explained the motivation: a 2.4-trillion-parameter model in BF16 approaches five terabytes of VRAM, and fitting a serving replica on one node could avoid the inter-node communication involved when that replica is split with tensor and pipeline parallelism.
The two-team Spark run and that frontier-scale work sit at different stages. One demonstrates that resident optimized models and separate team clusters can work together on a single host. The other describes where Minima aims to take model optimization on larger systems.