Efficient Local LLMon one DGX Spark.
Upload a document, ask in plain language, and get answers that cite their exact page — from an AI that runs entirely on one desk-side machine. Your documents never leave it.
Every layer, engineered by experts.
What feels effortless on the surface is a full AI infrastructure underneath — memory budgets, model serving, retrieval, governance, and experience, designed together as one system on one machine.
Effortless to use. Efficient by design.
Every layer beneath is researched, tuned, and optimized — to deliver maximum performance for you.
One machine runs everything.
No rack, no cloud, no integration project. The whole platform — models, index, and dashboards — lives in one quiet box on a desk.
- GB10 GRACE-BLACKWELL
- 128 GB UNIFIED MEMORY
- 1 MACHINE — EVERYTHING


The GB10 superchip: CPU and GPU on one die, sharing 128 GB of unified memory.
Intelligence, sized to the machine.
Two chat models and a Thai document reader share one memory pool — the platform decides what is awake, so users never think about it.
PreMize v0.1
The primary model for chat and summarization — deep answers over long documents.
PreMize Lite
A fast secondary model that sleeps when idle and wakes automatically when called.
Thai OCR
Reads scanned Thai paperwork page by page — contracts, handbooks, invoices.
Semantic search
Finds the right passage even when the question uses different words than the document.
Upload. Ask. Get cited answers.
No setup, no training, no manual. The platform reads, summarizes, and indexes every document automatically — then answers in plain language with the page number attached.
Measured, not promised.
Recorded on the production DGX Spark under concurrent load — not a projection.
Every answer cites its page.
Citations come from the document itself — deterministic “Ref page N” links to the exact source page. The model can never invent one. Trust is checkable in one click.
Full-time employees receive 12 days of annual leave per year.
Ref page 12Complete, and effortless.
Cited document chat
Ask across one document or your whole library — answers carry page references.
Built-in assistant
A floating guide that knows the platform and helps every user find their way.
Agent Studio
Compose agents visually; a test scorecard must pass before anything goes live.
Customer channels
Public web chat and LINE OA, answering customers with the same cited accuracy.
Open API
Standard-compatible API keys, self-served — connect your own tools in minutes.
Usage dashboards
Per-user metering and live system health, visible at a glance.
The product, as it ships.
Real screens from the running platform — Thai-first, cited, and observable end to end.

Cited document chat
Summary, conversation, and citations side by side — every claim carries a “Ref page N” chip that links straight to its source page.

Agent Studio
Compose an agent from tools, document scope, and policies on a visual canvas — and it must pass a test scorecard before it goes live.

Admin control tower
Tokens, requests, and cost — per user, per feature, per model. Fourteen days of the whole organization at a glance.

Chat with live web search
An optional governed web search adds fresh sources to general chat, each answer carrying its reference chips.

Model picker
Switch between the 35B and the fast 8B — capabilities shown up front.

Live GPU telemetry
Utilization, temperature, and power of the GB10, streamed in real time.

Self-serve API keys
A standard-compatible endpoint with per-user metering, one click away.

Your machine, in your pocket.
Hermes is an intelligent agent that lives on the DGX Spark and answers on Telegram. Wake models, check the GPU, ask about your documents — from your phone, anywhere, anytime. The conversation travels; the computing never leaves the machine.
- Chat on Telegram from any phone — anywhere, anytime
- Wake and sleep models, run prompts, check live GPU status
- Everything is processed locally on your DGX Spark
Powered by Hermes · Nous Research

A hand-picked toolset — each tool tuned and measured to run at its best on the appliance, and every action passes policy → approval → audit.
GPU load, temperature, power, and service health — live from the machine.
Wake and sleep models from chat; sensitive actions wait for your approval.
Ask about your documents — answers cite their exact page.
Live web results through the appliance's own search engine — no third-party API.
Send a photo in the chat; the machine reads it and explains what it sees.
Token usage by model, plus the approvals inbox — decide right in the chat.
The same machine. A different class of system.
A bare DGX Spark gives you excellent hardware — and months of hard engineering before it serves your business. PreMize ships with that engineering already done: we spent the research hours on memory budgets, serving behavior, retrieval, and governance so every layer runs at its optimum — resources, latency, and efficiency. Complex underneath, so it is simple on top.
DGX Spark only
DIYexcellent hardware — everything else is homework
- Partitioning 128 GB of unified memory (trial, error, OOM)
- Model serving, quantization, and tuning
- Thai OCR pipeline for scanned documents
- RAG with trustworthy page-level citations
- Governance, approvals, and audit
- Dashboards, metering, and monitoring
DGX Spark + PreMize
READYengineered, measured, ready on day one
- Every gigabyte of unified memory budgeted
- Serving tuned for the GB10 — with an auto sleep/wake steward
- Thai OCR built in, page by page
- Cited answers out of the box — “Ref page N”
- Policy → approval → audit, from day one
- Live dashboards, per-user metering, open API
Measured on the same DGX Spark — before and after our tuning
From the production load-test campaign — not a projection.
Your documents never leave the machine.
100% on-premises
Models, index, and documents live on your hardware.
Air-gap capable
Once installed, it runs with no internet at all.
PDPA-aligned
Data residency by construction, with a full audit trail.



