Enterprise deployments

Your models. Your infrastructure. Up to 4× fewer GPUs.

Minima optimises the models your enterprise already uses and delivers the optimised weights and mnma runtime for deployment in your environment.

Core offer

Optimisation built around your model and deployment.

01

Your chosen model

Optimise your chosen open-weight or internal model, subject to evaluation and access to the necessary artefacts.

02

Agreed hardware

Deliver the optimised model and compatible runtime for the agreed hardware.

03

Your infrastructure

Deploy within your infrastructure, including the existing supported on-premises and VPC paths.

04

Measured acceptance

Benchmark quality, throughput, memory and GPU requirements against your baseline and acceptance criteria.

What you receive

The validated deployment package.

  1. 01

    The agreed optimised model artefacts.

  2. 02

    The mnma serving runtime for the validated configuration.

  3. 03

    Deployment instructions and integration requirements.

  4. 04

    A comparison report against agreed quality and operating targets.

Featured deployment result

Qwen3.8-2.4T

Four NVIDIA B200 nodes → one NVIDIA B200 node

Four NVIDIA B200 nodes for the baseline become one NVIDIA B200 node with Minima.

4× smallerModel weights
3.5× smallerKV cache
2× higherTokens per second

Minima internal benchmark and featured deployment result. Achievable savings for a customer's model, workload and hardware are established through evaluation.

Workflow

From baseline to licensed deployment.

  1. 01Establish the baseline
  2. 02Optimise
  3. 03Validate quality and serving performance
  4. 04Deploy under the annual licence

Commercial model

Annual software licence, priced on your agreed GPU-hour usage.

We scope the model, hardware and workload with your team, validate the improvement, and agree the production licence.

LICENCE SCOPE

Agreed for the validated production deployment.

  • Model, hardware and workload scope
  • Validated improvement
  • Agreed production licence

Enterprise FAQ

Deployment, delivery and evaluation.

Where does the deployment run?

Deployment runs within the customer's infrastructure through the agreed supported on-premises or VPC path.

What model and runtime artefacts are supplied?

Minima supplies the agreed optimised model artefacts, the mnma serving runtime for the validated configuration, and deployment instructions and integration requirements.

How is quality evaluated?

Quality is benchmarked against the customer's baseline and agreed acceptance criteria alongside throughput, memory and GPU requirements.

How is GPU-hour licensing scoped?

Minima scopes the model, hardware and workload with the customer, validates the improvement, and prices the annual software licence on agreed GPU-hour usage.

Enterprise benchmark

Request an efficiency benchmark

Minima will scope your model, hardware and workload with your team and validate the improvement against the agreed baseline and acceptance criteria.

By submitting, you ask Minima to contact you about an Enterprise efficiency benchmark. See our Privacy Policy and Terms of Service.