Article
Minima leverages Qdrant to achieve 2.78x more agentic RAG tasks on one GPU
This is a joint benchmark of Qdrant retrieval and Minima-optimized inference with Qwen3.6-27B model on one RTX PRO 6000 Blackwell.
Read articleMinima Blog
Research notes and systems writing on efficient inference, agent infrastructure and the economics of useful intelligence.
Article
This is a joint benchmark of Qdrant retrieval and Minima-optimized inference with Qwen3.6-27B model on one RTX PRO 6000 Blackwell.
Read articleArticle
Minima solved the efficiency problem inside the inference worker: 53% more tokens from Qwen3.6-27B on Blackwell. Rafay Token Factory solved the operability problem around it: a governed, token-metered service.
Read articleBenchmark with us
Send us the model, hardware and workload you care about. We will propose a benchmark with explicit quality, throughput, memory and cost gates.
Benchmark your model