
Sovereign · The flagship
Your own AI. Fully offline. In your building.
A sealed, fully-offline AI appliance — open-weight models plus your own knowledge base, running entirely inside your facility. No cloud, no data egress, no prompt ever sent to anyone. A private drop-in alternative to ChatGPT and Claude for organizations that legally cannot send their data to a cloud AI.
Who it's for
Built for data that can't leave the building.
For: regulated organizations that cannot put their data in a cloud AI — healthcare handling PHI under HIPAA, state/local government and federal-civilian teams handling CUI, legal practices protecting privilege, financial services, and semiconductor or IP-heavy shops guarding trade secrets. Our beachhead is healthcare: no clearance required, a clear HIPAA driver, and a fast buy cycle.
Not for: anyone who needs frontier-model (GPT-5.x/Opus-tier) quality at cloud speed with no regulatory driver — cloud is cheaper and faster for them, and we'll say so. Not for classified/SCIF prime work either: that requires a facility clearance we don't hold yet.
The engagement · Level-2
How the intelligence actually moves into your building.
Owning the hardware is the destination. The Sovereignty Audit is the road.
It's a forward-deployed engagement — we come on-site, sit inside the work, and go department by department: sales, ops, finance, service. In each one we find where a person is doing what a system should, where your data is leaving the building to get answered, and where a cloud meter is quietly charging you rent on your own knowledge.
You walk away with two things you keep whether or not you build another thing with us:
- An ROI matrix — every opportunity we found, ranked by dollars freed against effort to capture. No hand-waving; the math is on the page.
- A Roadmap to Sovereignty — the ordered sequence that takes you from renting intelligence to owning it, rung by rung, ending at a private AI stack running on hardware you hold title to.
The two levels, honestly:
- Level 1 — Operations Audit ($2,500 flat). Two weeks, mostly remote, maps the leaks and hands you a prioritized automation plan. The front door. Yours to keep.
- Level 2 — Sovereignty Audit (priced per scope). On-site, forward-deployed, department by department. The full ROI matrix and the Roadmap to Sovereignty. Because scope varies enormously by organization, it's priced per engagement — book a scoping call and we'll size it to your operation.
The only outside number worth citing: MIT found 95% of enterprise AI pilots fail to deliver measurable value. They fail because they start with the model instead of the operation. We start with the operation. That's the whole difference.
Drakon Aegis
The operating layer your business runs on.
Drakon Aegis is the operating layer your business runs on — a live model of your operation you act from, with governed write-back into the systems you already run.
Most dashboards show you the problem. Aegis lets you fix it without leaving the screen.
We build a live model of your business — your jobs, customers, invoices, crews, inventory as real objects that know how they relate — and put a control surface on top of it. When something needs a decision, you make it right there: reroute the truck, approve the refund, escalate the late job. The action validates against your rules, writes back to your CRM, ERP, or books, and records who called it and when.
What's under it:
- A model of your business, not a pile of tables. Your operation as connected objects — a digital twin — so the screen understands this job, this crew, this customer, not raw rows.
- You act from the screen. Governed write-back into the systems you already run.
- Permissioned to the field. Who can see what, who can do what, down to the column and the action. An operator can flag; only a manager can close.
- Full audit trail. Every decision captured — what, who, when, on what data. Built for the regulated buyer who has to prove it later.
- Runs on AI you own. Wired into your private on-prem stack when you want it — your data never leaves the building to power the screen.
Aegis is custom by definition. We come in forward-deployed, learn how your business actually runs, map your stack, and build the model to fit. Nothing here is plug-and-play or overnighted — the road is the value, and it's what makes the result yours.
You see your entire business at every tier. What scales is how much of it Aegis actively runs.
- Aegis · Focused — from $15,000 build, then managed. See everything. Run your highest-leverage operation. The whole business on one screen from day one, and Aegis goes live acting where it hurts most: dispatch, billing, or intake. Full picture, first hand on the wheel.
- Aegis · Full — from $25,000 build, then managed. See everything. Run everything. Every core system wired for action, every department live on one object layer — your company run end to end from a single surface, by the people who own it.
- Aegis · Enterprise — priced to the engagement. A family of businesses, one control layer. Every entity in the portfolio on one screen, action org-wide, compliance-grade governance and audit, deployed on AI that never leaves your building. Sized on a scoping call, built to hold.
Own the layer your business runs on.
Start with the $2,500 Operations Audit — we map where Aegis pays for itself, then build it to fit.
The ladder
One owned box first. Present it as one ladder.
We sell a single owned box as the default — simpler, faster for interactive use, and honest. One continuity worth stating plainly: the Entry rung below is the same "On-Prem / Own-It · from $15,000" you saw on our home page. It is the entry rung of this ladder; the Sovereign appliance from $45K+ is the top of the same ladder, not a different product line.
| Tier | Hardware | Class | Runs | Interactive speed |
|---|---|---|---|---|
| Entry · On-Prem Own-It | from $15,000 | capable starter | 7B–13B class (sized to your budget/hardware) | interactive · scoped at quote |
| Standard | Mac Studio M4 Max 128GB | GPT-4-class | Llama 3.3 70B (Q4) | 15–22 tok/s (PROVEN) |
| Sovereign | Mac Studio M3 Ultra 256GB | GLM-5.2-class | 200B+ MoE (GLM-5.2 744B MoE, MIT, 1M ctx) | directional |
| Sovereign-Max | M3 Ultra 512GB (secondary-market) | near-frontier | largest open models | directional |
Every tier includes a local vector database (Qdrant/Milvus), OCR document ingestion, Open WebUI with role-based access control, and a compliance-documentation package (HIPAA / CMMC / ATO artifacts) — worth $20–40k standalone, and the real reason a buyer picks us over raw iron.
Clustering
Clustering is capacity, not speed — and we say so.
A single box is the default. When a client needs a model too big for one box (the 600B–1T class), we pool memory across boxes today — shipped hardware, not a promise. RDMA over Thunderbolt 5 shipped in macOS 26.2, and the M3 Ultra has TB5 on every port, so separate boxes read each other's unified memory at near-local latency. Independently proven: four M3 Ultras pooled ≈ 1.5TB of memory running Exo, with DeepSeek V3.1 671B producing 32.5 tok/s across the 4 nodes.
The honest caveats — which are also the moat, because we state them plainly:
- Capacity, not speed — pooling lets you run a model that won't fit one box; it does not make a model that fits go faster.
- Exo / MLX-Distributed only.
- 4-node ceiling today — no Thunderbolt 5 switch exists.
- New and not hardened — we validate the rig in-house before it ships.
- One box usually suffices; we lead with single-box and cluster only when the model demands it.
PROVEN · RDMA over Thunderbolt 5 + Exo, ~1.5TB pooled across 4 M3 UltraDIRECTIONAL · Tier 2/3 interactive tok/s — measured in-house before any quote
Pricing
Priced in the open. From $45,000.
The appliance flagship. Hardware BOM is roughly $4k (128GB) to $7k (256GB); Tier-3 512GB units are secondary-market-sourced — Apple discontinued the 512GB config in March 2026 and the 256GB in May 2026, and the Apple-direct ceiling is now 96GB. The balance is integration labor, the compliance-documentation package, sourcing & verification, and margin.
The operational SLA, billed quarterly: scheduled on-site visits to ingest new documents, patch CVEs, side-load updated open-weight models, and run thermal sweeps. Required — an air-gapped box can't pull automated updates.
Honest note: secondary-market hardware pricing is availability-indexed and confirmed at procurement, not a fixed quote — the market is thin and moves fast. The $45k entry and $4.5k/mo SLA are the public anchors; tier prices and the monthly SLA are directional until BOM and labor lock.
DIRECTIONAL · Tier prices firm up on BOM + labor lock
Objections
The questions a serious buyer asks.
"I can buy a Dell/NVIDIA server for that." You can — and raw iron still needs months of integration, model selection, RAG, hardening, and audit documentation before it does anything. Sovereign is operational in days with the compliance package included.
"Cloud AI is cheaper." For unregulated data, yes. For PHI/CUI, cloud is a compliance liability you can't buy back after a breach. Owning also beats metered GPU rental within roughly 12 months at steady use.
"Is it as good as ChatGPT?" The Standard tier is GPT-4-class; the Entry rung runs smaller starter models, and higher tiers run frontier-class open weights. The honest tradeoff: slightly behind the very latest cloud model, but it's yours and it's offline.
"Can you even buy a high-memory Mac Studio right now?" The shortage pulled Apple's 512/256GB configs; Apple sells 96GB direct. High-memory units are hard to find, not impossible — they still trade on the secondary market, we're the specialists who source and verify a real one, and we're M5-launch-ready.
Impact
What moves.
Own it
The summit of the ladder.
Own the intelligence your business runs on.
Start with the $2,500 Operations Audit — we map where a private, owned AI stack pays for itself, then build the ladder up to a fully sovereign appliance.
