A sealed steel vault with azure light bleeding from the seams — air-gapped, contained intelligence.

Drakon Forge · Owned hardware

Your own AI. Fully offline. In your building.

A sealed, fully-offline AI appliance — open-weight models plus your own knowledge base, running entirely inside your facility. No cloud, no data egress, no prompt ever sent to anyone. A private drop-in alternative to ChatGPT and Claude for organizations that legally cannot send their data to a cloud AI.

The thesis

Intelligence is a commodity. The edge is deployment.

Everyone can rent a frontier model. Almost no one owns the operation that turns it into margin — whose building it runs in, whose hardware holds it, whose hands are on the data. That's the whole game, and it's the one we play with you.

Who it's for

Built for data that can't leave the building.

For: regulated organizations that cannot put their data in a cloud AI — healthcare handling PHI under HIPAA, state/local government and federal-civilian teams handling CUI, legal practices protecting privilege, financial services, and semiconductor or IP-heavy shops guarding trade secrets. Our beachhead is healthcare: no clearance required, a clear HIPAA driver, and a fast buy cycle.

Not for: anyone who needs frontier-model (GPT-5.x/Opus-tier) quality at cloud speed with no regulatory driver — cloud is cheaper and faster for them, and we'll say so. Not for classified/SCIF prime work either: that requires a facility clearance we don't hold yet.

The ladder

One owned box first. Present it as one ladder.

We sell a single owned box as the default — simpler, faster for interactive use, and honest. One continuity worth stating plainly: the Entry rung below is the same "Drakon Forge · from $15,000" you saw on our home page. It is the entry rung of this ladder; the Sovereign appliance from $45K+ is the top of the same ladder, not a different product line.

TierHardwareClassRunsInteractive speed
Entry · On-Prem Own-Itfrom $15,000capable starter7B–13B class (sized to your budget/hardware)interactive · scoped at quote
StandardMac Studio M4 Max 128GBGPT-4-classLlama 3.3 70B (Q4)15–22 tok/s (PROVEN)
SovereignMac Studio M3 Ultra 256GBGLM-5.2-class200B+ MoE (GLM-5.2 744B MoE, MIT, 1M ctx)directional
Sovereign-MaxM3 Ultra 512GB (secondary-market)near-frontierlargest open modelsdirectional

Every tier includes a local vector database (Qdrant/Milvus), OCR document ingestion, Open WebUI with role-based access control, and a compliance-documentation package (HIPAA / CMMC / ATO artifacts) — worth $20–40k standalone, and the real reason a buyer picks us over raw iron.

Clustering

Clustering is capacity, not speed — and we say so.

A single box is the default. When a client needs a model too big for one box (the 600B–1T class), we pool memory across boxes today — shipped hardware, not a promise. RDMA over Thunderbolt 5 shipped in macOS 26.2, and the M3 Ultra has TB5 on every port, so separate boxes read each other's unified memory at near-local latency. Independently proven: four M3 Ultras pooled ≈ 1.5TB of memory running Exo, with DeepSeek V3.1 671B producing 32.5 tok/s across the 4 nodes.

The honest caveats — which are also the moat, because we state them plainly:

  1. Capacity, not speed — pooling lets you run a model that won't fit one box; it does not make a model that fits go faster.
  2. Exo / MLX-Distributed only.
  3. 4-node ceiling today — no Thunderbolt 5 switch exists.
  4. New and not hardened — we validate the rig in-house before it ships.
  5. One box usually suffices; we lead with single-box and cluster only when the model demands it.

PROVEN · RDMA over Thunderbolt 5 + Exo, ~1.5TB pooled across 4 M3 UltraDIRECTIONAL · Tier 2/3 interactive tok/s — measured in-house before any quote

Pricing

Priced in the open. From $45,000.

from $45,000 one-time

The appliance flagship. Hardware BOM is roughly $4k (128GB) to $7k (256GB); Tier-3 512GB units are secondary-market-sourced — Apple discontinued the 512GB config in March 2026 and the 256GB in May 2026, and the Apple-direct ceiling is now 96GB. The balance is integration labor, the compliance-documentation package, sourcing & verification, and margin.

from $4,500/mo SLA

The operational SLA, billed quarterly: scheduled on-site visits to ingest new documents, patch CVEs, side-load updated open-weight models, and run thermal sweeps. Required — an air-gapped box can't pull automated updates.

Honest note: secondary-market hardware pricing is availability-indexed and confirmed at procurement, not a fixed quote — the market is thin and moves fast. The $45k entry and $4.5k/mo SLA are the public anchors; tier prices and the monthly SLA are directional until BOM and labor lock.

DIRECTIONAL · Tier prices firm up on BOM + labor lock

Objections

The questions a serious buyer asks.

"I can buy a Dell/NVIDIA server for that." You can — and raw iron still needs months of integration, model selection, RAG, hardening, and audit documentation before it does anything. Sovereign is operational in days with the compliance package included.

"Cloud AI is cheaper." For unregulated data, yes. For PHI/CUI, cloud is a compliance liability you can't buy back after a breach. Owning also beats metered GPU rental within roughly 12 months at steady use.

"Is it as good as ChatGPT?" The Standard tier is GPT-4-class; the Entry rung runs smaller starter models, and higher tiers run frontier-class open weights. The honest tradeoff: slightly behind the very latest cloud model, but it's yours and it's offline.

"Can you even buy a high-memory Mac Studio right now?" The shortage pulled Apple's 512/256GB configs; Apple sells 96GB direct. High-memory units are hard to find, not impossible — they still trade on the secondary market, we're the specialists who source and verify a real one, and we're M5-launch-ready.

Impact

What moves.

0 bytesData egress to third-party AI, eliminated.
15–22 tok/sInteractive inference · Standard tier · 70B Q4 (PROVEN).

Own it

The summit of the ladder.

Forge runs the hardware; Drakon Aegis is the operating layer that runs on top of it.

Own the intelligence your business runs on.

Start with the $2,500 Operations Audit — we map where a private, owned AI stack pays for itself, then build the ladder up to a fully sovereign appliance.