Ikioma Mini
The desk-sized entry point — two linked nodes, one private stack.
The recommended tier — 384 GB VRAM, four Blackwell GPUs, one integrated stack.
Hardware, delivered and deployed
Your data never leaves the floor. Air-gappable, no egress, no telemetry — the trust boundary has a serial number you can point at in an audit.
Own the machine and agents run around the clock. No per-token meter, no overage invoices — you pay for electricity, not tokens.
Hardware is a capital asset, not a subscription line. Weights, tunes, and stack stay yours — perpetually, with trade-in credit on next-gen silicon.
Open-weight foundation models, post-trained on the categories of work your team actually does — document understanding, summarisation, structured extraction, code, agentic tool use. We benchmark against your task suite, not generic leaderboards — and the weights are yours, perpetually.
Own the machine and your agents work all night — no rate limits, no per-token meter, no "try again later". Coding loops and long-horizon tasks run until they're done, and 128k tokens of context keeps them coherent start to finish.
The runtime drops in as an OpenAI- and Anthropic-compatible API, so your existing tools keep working. We handle model updates, security patches, and performance tuning; updates are signed and reversible, and a named engineer owns your account.
Private inference is CapEx — it goes on the balance sheet, not the expense line. Most customers break even versus equivalent cloud spend within 4–7 months on typical agentic workloads; after that, the machine keeps compounding.
Assembled in Yokohama and Eindhoven. Boards burn in for 96 hours before they leave the floor. Every system carries a serial, a calibration sheet, and a name etched on the chassis — because the people who built it stand behind it for a decade, not until next quarter's earnings call.
Preinstalled and pre-validated: the ikioma inference runtime, your choice of optimized open models, and the production middle layer (audit log, RBAC, evals) configured out of the box.
| Model | Params | Context | tok/s | Latency | Notes |
|---|---|---|---|---|---|
| Qwen 3 72B | 72B | 128k | — | — | TBD — burn-in benchmark |
| Llama 3.3 70B | 70B | 128k | — | — | TBD — burn-in benchmark |
Loading 3D view…
Four Blackwell GPUs and the full cooling array. Drag to orbit, run the explode slider to lift the assembly apart layer by layer, X-ray the shell to see inside — or switch the configuration to the open frame or the GPU tower.
Illustrative 3D approximation — not manufacturing CAD. Schematic dimensions only.
Mini for the desk, Ikioma for the team, Pro for the fleet. Every version ships the same model and inference stack — the scale is what changes.