Minima API

Faster tokens for the models you want.

Minima provides an OpenAI-compatible API for optimized inference, giving teams a familiar interface to models served through the Minima runtime.

API and value

Optimized inference through an interface your stack already knows.

Use an OpenAI-compatible API to access supported models through Minima's optimized serving layer.

Interface

Compatible interface

Use an OpenAI-compatible API and request format.

Serving

Optimized inference

Access supported models served with Minima optimization.

Models

Growing model set

Choose from the listed model families, with more models coming.

Supported models

Model availability for the workloads you choose.

Minima API model availability includes the Qwen3.8 family, GLM5.3 Flash, Gemma 4-31B and Trinity, with more models coming.

  • Qwen3.8 family
  • GLM5.3 Flash
  • Gemma 4-31B
  • Trinity
  • More models coming

Performance and commercial model

More throughput, straightforward access.

Minima measures API throughput in tokens per second and offers both usage-based and committed access models.

Minima internal measurement

50%higher TPS
1.5×vs most OpenRouter providers

Minima internal measurement. This result has not been independently verified and is not a universal performance guarantee. Actual performance depends on the model, workload, provider configuration and serving conditions.

Commercial model

Choose the access model that fits your usage.

  • Pay as you go Usage-based pricing for supported models.
  • Committed usage 10% off applicable pay-as-you-go rates for committed usage.

Contact Minima to agree volume and terms.

API waitlist

Join the Minima API waitlist.

Register interest in OpenAI-compatible access to Minima-optimized models.

By submitting, you ask Minima to contact you about API access. See our Privacy Policy and Terms of Service.