← All systems
Made to order · 2–8 weeks

Ikioma

The recommended tier — 384 GB VRAM, four Blackwell GPUs, one integrated stack.

From €40,000

Hardware, delivered and deployed

Support & optimization €13,000/year
Year 1 total €53,000
Configure yours →

Or book a deployment call directly →

[ why own it ]
01 · physics

Private by physics.

Your data never leaves the floor. Air-gappable, no egress, no telemetry — the trust boundary has a serial number you can point at in an audit.

02 · metering

Pay once, infer forever.

Own the machine and agents run around the clock. No per-token meter, no overage invoices — you pay for electricity, not tokens.

03 · ownership

Yours to keep.

Hardware is a capital asset, not a subscription line. Weights, tunes, and stack stay yours — perpetually, with trade-in credit on next-gen silicon.

01 · your work

Models that know your work.

Open-weight foundation models, post-trained on the categories of work your team actually does — document understanding, summarisation, structured extraction, code, agentic tool use. We benchmark against your task suite, not generic leaderboards — and the weights are yours, perpetually.

72B
class of open models
128ktok
context — long documents, long tasks
100%
of weights owned by you
02 · autonomy

Agents run around the clock.

Own the machine and your agents work all night — no rate limits, no per-token meter, no "try again later". Coding loops and long-horizon tasks run until they're done, and 128k tokens of context keeps them coherent start to finish.

24/7
unmetered inference
128ktok
context for long tasks
∞
no rate limits, no overages
03 · operations

You don't need to become an ML team.

The runtime drops in as an OpenAI- and Anthropic-compatible API, so your existing tools keep working. We handle model updates, security patches, and performance tuning; updates are signed and reversible, and a named engineer owns your account.

/v1
drop-in — OpenAI / Anthropic compatible
1binary
self-hosted runtime
1owner
named engineer per account
04 · economics

A capital asset, not a subscription.

Private inference is CapEx — it goes on the balance sheet, not the expense line. Most customers break even versus equivalent cloud spend within 4–7 months on typical agentic workloads; after that, the machine keeps compounding.

4–7mo
typical break-even vs cloud
10yr
security updates, in writing
€0
per token, after purchase
勢
[ heritage ]

Built in the lineage of instruments, not platforms.

Assembled in Yokohama and Eindhoven. Boards burn in for 96 hours before they leave the floor. Every system carries a serial, a calibration sheet, and a name etched on the chassis — because the people who built it stand behind it for a decade, not until next quarter's earnings call.

96h
burn-in per system
before it leaves the factory floor
2
assembly sites
Yokohama · Eindhoven
10yr
security update commitment
in writing
[ the system ]

Hardware spec

GPU 4× RTX Pro 6000 Blackwell
VRAM 384 GB
CPU 32-core AMD Genoa
RAM 192 GB
Storage 4 TB high-speed RAID + 1 TB boot
PSU 2× 1600W
Dimensions 19″ W × 21″ H × 16.25″ D
Weight 90 lbs
lead_time Made to order · 2–8 weeks

Model & inference stack Included

Preinstalled and pre-validated: the ikioma inference runtime, your choice of optimized open models, and the production middle layer (audit log, RBAC, evals) configured out of the box.

  • ✓OpenAI / Anthropic API compatibility
  • ✓Sandboxed tool execution
  • ✓Signed audit log
  • ✓Runs on-premises — your hardware, your premises
Model Params Context tok/s Latency Notes
Qwen 3 72B 72B 128k — — TBD — burn-in benchmark
Llama 3.3 70B 70B 128k — — TBD — burn-in benchmark

Loading 3D view…

Configuration
Drag to rotate · Scroll to zoom
the hardware

The workhorse, up close.

Four Blackwell GPUs and the full cooling array. Drag to orbit, run the explode slider to lift the assembly apart layer by layer, X-ray the shell to see inside — or switch the configuration to the open frame or the GPU tower.

Illustrative 3D approximation — not manufacturing CAD. Schematic dimensions only.