All systems live · serving from our own racks

AI cloud on
our own metal

We build and operate the data centers behind this API. OpenAI-compatible inference on hardware we rack and tune ourselves, priced per token in prepaid credits. GPUs by the hour follow as the fleet comes online.

100 crfree at signup, no card
/v1OpenAI-compatible API
256Kcontext window today
0third-party inference providers
01 Products

Two ways onto our hardware

Live

Serverless Inference

Open models behind Chat Completions, Responses, and Anthropic Messages, plus a streaming chat UI. Per-token pricing in credits, published on every model. Streams over SSE.

Browse models and prices
Waitlist

GPU Cloud

Dedicated GPUs by the hour for training, fine-tuning, and self-managed inference, on machines we rack ourselves. The fleet is in commissioning now.

See hardware and join the waitlist
02 Models

Live catalog, live prices

The same table billing reads from. 1 credit = $0.01, deducted per token.

ModelContextInput / 1M tokensOutput / 1M tokens
Qwen3 Coder 30B A3Bchat-default64K10 cr ($0.10)30 cr ($0.30)
Qwen 3.6qwen3.6256K40 cr ($0.40)120 cr ($1.20)
03 Infrastructure

We run the racks

Most AI APIs are a margin on someone else's cloud. Ours is the cloud.

An engineer cabling servers in a rack hands on the metal

Our racks, our rules

Inference runs in our own data centers, on machines we bought, racked, and tuned ourselves. Nobody sits between you and the hardware.

Isolation in the database

Tenant separation is row-level security in Postgres. The database enforces it on every query, beneath any application code.

Honest capacity

The admission queue is sized to measured throughput. Under load you see your position in the queue or an explicit timeout. Quality never degrades silently.

Published numbers

Benchmarks come from real runs on the machines that serve you. Every model card prices input and output per 1M tokens, to the credit.

// platform
endpoint api.hakk.ai/v1
compat OpenAI + Anthropic · SSE
models 2 live · growing
context 256K per request
billing prepaid credits · per token
status in production

Served from data centers we own and operate. The GPU Cloud page shows what is coming next.

04 How it works

From signup to first token in minutes

01

Create an account

You start with 100 free credits and no card. Your account is an organization from day one.

02

Get a key or open chat

Generate an API key and point your editor at the contract on the docs page. The streaming chat workspace is ready the same minute.

03

Pay per token

Credits are deducted at the model's published rate. Top up in fixed packages, and every movement lands in a ledger you can read.

05 Credits

Prepaid credits, plain prices

1 credit is one cent. Credits never expire, and there is no subscription.

Starter$101,000 creditsEnough credits to evaluate every model against your real workload.Request access
Builder$252,500 credits+125 bonusFor daily driving an editor or agent against the API.Request access
Scale$10010,000 credits+1,000 bonusTeam workloads and batch jobs. Contact us past this size.Request access
06 FAQ

Common questions

hakk.ai is an infrastructure operator. We own data centers and the machines in them, and inference runs on that hardware directly. The fleet is growing, with more machines in commissioning.

Put your tokens on our metal.

We onboard in small batches and set up every account personally. Request access and start with 100 free credits.