A sealed steel vault with azure light bleeding from the seams — air-gapped, contained intelligence.

Drakon Forge · Owned hardware

Your own AI. Fully offline. In your building.

A sealed, fully-offline AI appliance — open-weight models running entirely inside your facility, on hardware you own, with your raw company data loaded into a private store you control. No cloud, no data egress, no prompt ever sent to anyone. A private drop-in alternative to ChatGPT and Claude. Add the advanced reasoning brain — the part that understands and governs your knowledge — with Drakon Cipher.

The thesis

Intelligence is a commodity. The edge is deployment.

Everyone can rent a frontier model. Almost no one owns the operation that turns it into margin — whose building it runs in, whose hardware holds it, whose hands are on the data. That's the whole game, and it's the one we play with you.

Who it's for

Built for data that can't leave the building.

For: regulated organizations that cannot put their data in a cloud AI — healthcare handling PHI under HIPAA, state/local government and federal-civilian teams handling CUI, legal practices protecting privilege, financial services, and semiconductor or IP-heavy shops guarding trade secrets. Our beachhead is healthcare: no clearance required, a clear HIPAA driver, and a fast buy cycle.

Not for: anyone who needs frontier-model (GPT-5.x/Opus-tier) quality at cloud speed with no regulatory driver — cloud is cheaper and faster for them, and we'll say so. Not for classified/SCIF prime work either: that requires a facility clearance we don't hold yet.

The ladder

One owned box first. Present it as one ladder.

We sell a single owned box as the default — simpler, faster for interactive use, and honest. One continuity worth stating plainly: the Entry rung below is the same "Drakon Forge · from $15,000" you saw on our home page. It is the entry rung of this ladder; the Apple-silicon rungs climb to a 512GB appliance from $45K+, and the Enterprise rung (NVIDIA GB300) tops the same ladder — one product line, not several.

TierCapacityClassRunsInteractive speed
Entry · On-Prem Own-It96GB unified memorycapable starter7B–13B class (sized to your budget)interactive · scoped at quote
Standard128GB unified memoryGPT-4-classLlama 3.3 70B (Q4)15–22 tok/s (PROVEN)
Sovereign256GB unified memoryGLM-5.2-class200B+ MoE (GLM-5.2 744B MoE, MIT, 1M ctx)directional
Sovereign-Max512GB unified memorynear-frontierlargest open modelsdirectional
Enterprise748GB coherent (NVIDIA GB300)frontier-class405B on one box · 200–300 concurrent usersclusters over 800 Gb/s fabric

Every tier includes a local vector database (Qdrant/Milvus), OCR document ingestion, Open WebUI with role-based access control, and a compliance-documentation package (HIPAA / CMMC / ATO artifacts) — worth $20–40k standalone, and the real reason a buyer picks us over raw iron.

Turning that loaded data into a governed brain that reasons over it — semantic understanding, freshness governance, fine-tuning — is Drakon Cipher, sold on its own or bundled with Forge.

Clustering

Two lineages, one honest story: clustering is capacity, not speed.

A single box is the default on every rung. When a client needs a model too big for one box (the 600B–1T class), we pool across boxes — and how we pool depends on which lineage you're on.

Apple-silicon rungs (96–512GB). RDMA over Thunderbolt 5 shipped in macOS 26.2, and Apple silicon has TB5 on every port, so separate boxes read each other's unified memory at near-local latency. Independently proven: four M3 Ultras pooled ≈ 1.5TB of memory running Exo, with DeepSeek V3.1 671B producing 32.5 tok/s across the 4 nodes. Honest caveats — which are also the moat, because we state them plainly: it's Exo / MLX-Distributed only; there's a 4-node ceiling today (no Thunderbolt 5 switch exists yet); and it's new and not hardened, so we validate every rig in-house before it ships.

Enterprise rung (NVIDIA GB300). A different, datacenter-grade fabric: each GB300 is DGX Station-class with 748GB coherent memory (252GB HBM3e @ 7.1 TB/s + 496GB LPDDR5X) and an NVLink-C2C link at 900 GB/s between its own Grace CPU and Blackwell GPU. Units cluster over an 800 Gb/s ConnectX-8 SuperNIC — a hardened, production networking stack, not an experimental pooling rig. One box already runs a 405B model or 200–300 concurrent users; you add units when you need more capacity.

The through-line across both: clustering buys capacity, not speed. Pooling lets you run a model that won't fit one box; it does not make a model that already fits run faster. One box usually suffices — we lead with single-box and cluster only when the model or the user count demands it.

PROVEN · RDMA over Thunderbolt 5 + Exo, ~1.5TB pooled across 4 M3 UltraPROVEN · NVIDIA GB300 · 800 Gb/s ConnectX-8 fabric, 748GB coherent per unitDIRECTIONAL · Tier 2/3 interactive tok/s — measured in-house before any quote

Pricing

Priced in the open. From $15,000 + $2,500/mo.

One ladder, priced end to end. The Entry rung starts at $15,000 + $2,500/mo managed and climbs to the Sovereign-Max 512GB tier at $45,000 + $7,500/mo. Across those four rungs the hardware BOM runs roughly $4k (128GB) to $7k (256GB); the 512GB tier is secondary-market-sourced — Apple discontinued the 512GB config in March 2026 and the 256GB in May 2026, and the Apple-direct ceiling is now 96GB. The balance is integration labor, the compliance-documentation package, sourcing & verification, and margin.

The Enterprise rung — NVIDIA GB300, 748GB coherent memory, frontier-class open models (Llama-3.1-405B-class) behind your own firewall — is a different hardware lineage and is priced by scope, on a custom quote, because the box is configured to your model and user count. There is no public number by design.

Honest note: secondary-market hardware pricing is availability-indexed and confirmed at procurement, not a fixed quote — the market is thin and moves fast. Tier prices and the monthly managed retainer are directional until BOM and labor lock.

DIRECTIONAL · Tier prices firm up on BOM + labor lock

Objections

The questions a serious buyer asks.

"I can buy a Dell/NVIDIA server for that." You can — and raw iron still needs months of integration, model selection, RAG, hardening, and audit documentation before it does anything. Sovereign is operational in days with the compliance package included.

"Cloud AI is cheaper." For unregulated data, yes. For PHI/CUI, cloud is a compliance liability you can't buy back after a breach. Owning also beats metered GPU rental within roughly 12 months at steady use.

"Is it as good as ChatGPT?" The Standard tier is GPT-4-class; the Entry rung runs smaller starter models, and higher tiers run frontier-class open weights. The honest tradeoff: slightly behind the very latest cloud model, but it's yours and it's offline.

"Can you even buy a high-memory Mac Studio right now?" The shortage pulled Apple's 512/256GB configs; Apple sells 96GB direct. High-memory units are hard to find, not impossible — they still trade on the secondary market, we're the specialists who source and verify a real one, and we're M5-launch-ready.

Impact

What moves.

0 bytesData egress to third-party AI, eliminated.
15–22 tok/sInteractive inference · Standard tier · 70B Q4 (PROVEN).

Bundle

À la carte, or build the stack and save.

Each stands on its own — buy exactly what you need. Or combine them and the quote comes down:

  • Forge + Cipher — 15% off the combined quote. The vault plus the mind: a private AI that truly understands your business, on hardware you own.
  • Forge + Cipher + Aegis — 20% off the combined quote. The full sovereign stack — vault, mind, and the operating layer your business runs from.

Discounts apply to the combined custom quote, not to published floors. Scoped from what your Operations Audit finds.

Own it

The summit of the ladder.

Forge is the vault — your data, owned, in your building; Drakon Cipher is the mind on top that reasons over it; Drakon Aegis is the operating layer. Own the vault, add the mind, run the business — each sold on its own, or bundled to save.

Own the intelligence your business runs on.

Start with the $2,500 Operations Audit — we map where a private, owned AI stack pays for itself, then build the ladder up to a fully sovereign appliance.